Site Reliability Engineer – Day Shift

🔥 2 minutes ago

🇺🇸 United States – Remote

💵 $104k - $166k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Peraton

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Peraton is a mission-focused enterprise that supports national security initiatives through advanced IT and cyber services. They provide capabilities in areas such as cyber defense, cloud operations, engineering, and intelligence. With a commitment to solving complex challenges, Peraton integrates data-driven technologies to ensure mission success for their military and government clients.

📋 Description

• Operate and maintain production infrastructure services and applications, ensuring availability, reliability, performance, security, and operational health • Monitor services and applications using SLIs, SLOs, dashboards, alerts, and observability tools; improve detection, diagnosis, and resolution of operational issues • Partner with application teams to define observability requirements and implement metrics, logs, traces, dashboards, and alerts • Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions • Execute application and infrastructure releases through deployment pipelines, including staging and production promotion, validation, rollback, and release troubleshooting • Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes • Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing • Identify and address reliability risks and operational technical debt using reliability metrics, incident trends, capacity data, and service health indicators • Automate operational activities using an everything-as-code approach • Collaborate with platform engineering and application teams to identify operational requirements and improve infrastructure building blocks, reliability, and operability

🎯 Requirements

• Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance • Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation • Proficient in Linux and Windows Server administration • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction • Scripting/automation proficiency in Python, Bash, PowerShell, or Go • Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53) • Preferred certifications: AWS Solutions Architect, AWS DevOps Engineer, AWS SysOps, Red Hat Certified Specialist in ROSA, Red Hat Certified System Administrator in OpenShift, Azure Administrator Associate, GCP Associate Cloud Engineer, Dynatrace Associate, Datadog Log Management Fundamentals, GitLab CI/CD Associate, Certified Jenkins Engineer (CJE), or Terraform Associate

🏖️ Benefits

• Potential eligibility for overtime • Potential eligibility for shift differential • Potential eligibility for a discretionary bonus

Apply Now

Similar Jobs

🔥 3 hours ago

Gifthealth

501 - 1000

🏥 Healthcare

📦 Logistics

💼 Consulting

DevSecOps Engineer embedding automated security across Gifthealth’s prescription healthcare platform. Building CI/CD, application, infrastructure, container, and Kubernetes security controls.

🇺🇸 United States – Remote

💵 $115k - $165k / year

💰 $40M Private Equity Round - GiftHealth on 2023-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 6 hours ago

Nagarro

10,000+ employees

💼 Consulting

📣 Marketing

🏥 Healthcare

Staff SRE Engineer building scalable AWS cloud platforms and reusable TypeScript infrastructure components. Advancing reliability, automation, observability, and DevSecOps across distributed microservices teams at Nagarro.

🔥 7 hours ago

Replit

51 - 200

🤖 Artificial Intelligence

🤝 B2B

Senior Site Reliability Engineer ensuring Replit’s reliable, scalable infrastructure serving millions of developers. Automating operations, observability, incident response, and performance optimization.

🔥 13 hours ago

TherapyNotes, LLC

51 - 200

💼 Consulting

⚖️ Legal

🏥 Healthcare

Site Reliability Engineer improving reliability, resilience, and observability for TherapyNotes’ behavioral health practice management and EHR SaaS platform. Automating operations across cloud infrastructure and 24×7 production services.

🔥 16 hours ago

Tandem Diabetes Care

1001 - 5000

🏥 Healthcare

🏭 Manufacturing

🔧 Hardware

Principal SRE ensuring reliable cloud infrastructure for Tandem Diabetes Care’s insulin technology. Leading incident response, automation, disaster recovery, and compliance across distributed teams.