Search Remote Jobs

Senior Cloud Site Reliability Engineer, SRE

Job not on LinkedIn

đŸ”„ 40 minutes ago

đŸ‡ș🇾 United States – Remote

đŸ’” $104k - $166k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🩅 H1B Visa Sponsor

infoinfo

đŸ‘» Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Peraton

Peraton

10,000+ employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Consulting ‱ Healthcare ‱ Logistics

Peraton is a mission-focused enterprise that supports national security initiatives through advanced IT and cyber services. They provide capabilities in areas such as cyber defense, cloud operations, engineering, and intelligence. With a commitment to solving complex challenges, Peraton integrates data-driven technologies to ensure mission success for their military and government clients.

📋 Description

‱ Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments ‱ Build automation scripts, APIs, and utilities in Python to reduce toil and improve platform reliability ‱ Implement observability and monitoring solutions using Grafana and AWS CloudWatch ‱ Build and optimize Terraform Infrastructure as Code for AWS resources and SRE solutions ‱ Develop CI/CD pipelines and automated testing ‱ Define SRE standards, best practices, guidelines, and metrics such as SLIs and SLOs ‱ Apply version control, code reviews, test-driven development, and documentation practices ‱ Participate in incident management and an on-call rotation ‱ Provide technical support for SRE tools and troubleshoot production issues ‱ Collaborate with teams to reduce incident recurrence through proactive detection and pattern analysis ‱ Stay current with AWS services, SRE methodologies, and cloud-native development technologies ‱ Collaborate with cross-functional teams in Agile and Scaled Agile frameworks ‱ Produce clear, blameless postmortems with actionable items and documented failure scenarios

🎯 Requirements

‱ Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level Clearance ‱ Bachelor's Degree and 8 years of experience, or a High School diploma or equivalent and 12 years of experience ‱ 5+ years of advanced Python development experience building enterprise-grade, highly available tools, APIs, and utilities for AWS ‱ 7+ years of software development experience focused on reliability and platform engineering ‱ 3+ years of hands-on experience developing solutions in AWS environments ‱ Deep understanding of AWS services including EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, and Step Functions ‱ Experience with AWS resource cost optimization ‱ 3+ years applying SRE principles including observability, toil automation, SLIs/SLOs, and reliability engineering ‱ Expert-level proficiency with Terraform IaC, including module development and state management ‱ Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices ‱ Experience with Grafana, AWS CloudWatch, and AWS Canary ‱ Experience defining, implementing, and managing SLOs/SLIs and error budgets ‱ Familiarity with conducting RCAs and producing postmortem documentation ‱ Working experience in Agile and Scaled Agile environments ‱ Familiarity with ITSM processes, resilience testing, and chaos engineering practices

đŸ–ïž Benefits

‱ Employees may be eligible for overtime ‱ Employees may be eligible for shift differential ‱ Employees may be eligible for a discretionary bonus ‱ Equal opportunity employer, including disability and protected veterans

Apply Now

Similar Jobs

đŸ”„ 8 hours ago

The SSI Group

11 - 50

🎯 Recruiter

đŸ‘„ HR Tech

đŸ€ B2B

DevOps Process Analyst optimizing development and operations workflows for SSI Group, a healthcare software company. Improving processes, reporting, risk management, and cross-functional delivery.

đŸ”„ 10 hours ago

CentralReach

201 - 500

đŸ„ Healthcare

📚 Education

Senior SRE improving reliability, observability, and cloud infrastructure for CentralReach’s autism and IDD care software platforms. Driving SLO adoption, incident response, automation, and production performance.

đŸ”„ 10 hours ago

CentralReach

201 - 500

đŸ„ Healthcare

📚 Education

Site Reliability Engineer operating AWS and Snowflake infrastructure for CentralReach’s autism and IDD care software. Improving reliability, connectivity, deployments, and incident response across data platforms.

đŸ‡ș🇾 United States – Remote

đŸ’” $135k - $160k / year

💰 Private equity on 2018-03

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ”„ 10 hours ago

Katmai Government Services

1001 - 5000

đŸ›ïž Government

📋 Compliance

⚕ Healthcare Insurance

DevOps Engineer II automating CI/CD, containers, and infrastructure for Navy readiness and training systems. Supporting DoD cybersecurity compliance across development, test, and production environments.

đŸ”„ 11 hours ago

Laravel

11 - 50

☁ SaaS

🌐 Web 3

🏱 Enterprise

Senior Site Reliability Engineer building Laravel’s multi-region Kubernetes infrastructure and observability systems. Establishing SRE practices, SLOs, and automation across global developer products.