Site Reliability Engineer – Night Shift

🔥 14 hours ago

🇺🇸 United States – Remote

💵 $104k - $166k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Peraton

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Peraton is a mission-focused enterprise that supports national security initiatives through advanced IT and cyber services. They provide capabilities in areas such as cyber defense, cloud operations, engineering, and intelligence. With a commitment to solving complex challenges, Peraton integrates data-driven technologies to ensure mission success for their military and government clients.

📋 Description

• Operate and maintain production infrastructure services and applications, ensuring availability, reliability, performance, security, and operational health • Monitor services and applications using SLIs, SLOs, dashboards, alerts, and observability tools • Define application observability requirements with application teams and implement metrics, logs, traces, dashboards, and alerts • Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions • Execute application and infrastructure releases through deployment pipelines, including staging and production promotion, validation, rollback, and release troubleshooting • Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes • Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing • Identify and address reliability risks and operational technical debt using reliability metrics, incident trends, capacity data, and service health indicators • Automate operational activities using an everything-as-code approach • Collaborate with platform engineering and application teams to identify operational requirements and improve environment reliability and operability

🎯 Requirements

• Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance • Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation • Proficient in Linux and Windows Server administration • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction • Scripting/automation proficiency in Python, Bash, PowerShell, or Go • Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53) • Preferred: AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification • Preferred: Red Hat Certified Specialist in ROSA or Red Hat Certified System Administrator in OpenShift • Preferred: Azure Administrator Associate or GCP Associate Cloud Engineer certification • Preferred: Dynatrace Associate or Datadog Log Management Fundamentals certification • Preferred: GitLab CI/CD Associate certification or Certified Jenkins Engineer (CJE) • Preferred: Terraform Associate certification

🏖️ Benefits

• Overtime eligibility may apply • Shift differential may apply • Discretionary bonus may apply

Apply Now

Similar Jobs

🔥 15 hours ago

Rackner

11 - 50

💼 Consulting

🎖️ Defense

🤖 Artificial Intelligence

R&D DevSecOps Engineer building secure AI-enabled mission software and DevSecOps pipelines. Developing backend services, coding agents, and compliant delivery workflows for U.S. government defense customers.

🔥 20 hours ago

AXON Networks

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Site Reliability Engineer improving reliability across AXON Networks’ AI-driven ISP cloud platform and high-speed routers. Automating NOC operations, observability, incident response and device-management recovery.

🕒 2 days ago

NetCov

201 - 500

🔒 Cybersecurity

🤝 B2B

💼 Consulting

AI Deployment Engineer deploying secure AI solutions across Hatz AI, Microsoft Co-Pilot, and Anthropic Claude. Supporting NetCov’s IT and cybersecurity services through data governance, troubleshooting, and customer implementations.

🇺🇸 United States – Remote

💵 $90k - $150k / year

💰 Private equity on 2022-11

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

NetCov

201 - 500

💼 Consulting

🏥 Healthcare

🔒 Cybersecurity

AI Deployment Engineer deploying secure AI solutions across Hatz AI, Microsoft Co-Pilot, and Anthropic Claude. Supporting NetCov’s customer IT and cybersecurity environments through governance, troubleshooting, and integration.

🇺🇸 United States – Remote

💵 $90k - $150k / year

💰 Private equity on 2022-11

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

Avalon Healthcare Solutions

51 - 200

🏥 Healthcare

☁️ SaaS

⚕️ Healthcare Insurance

Site Reliability Engineer securing and scaling AWS infrastructure for Avalon Healthcare Solutions’ diagnostic intelligence platform. Automating cloud operations, observability, and network security with Terraform and Palo Alto.