Senior Site Reliability Engineer

🔥 1 minute ago

🇺🇸 United States – Remote

💵 $210k - $275k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 21%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Replit

Replit

51 - 200 employees

🤖 Artificial Intelligence

🤝 B2B

Artificial Intelligence • B2B

Replit is a multiplayer computer platform designed for creating and sharing software collaboratively from anywhere in the world and on any device. The platform offers a feature called Generate Code, which utilizes AI to assist users in writing code, significantly streamlining the development process.

📋 Description

• Join Replit’s Site Reliability Engineering team to ensure the reliability, scalability, and performance of infrastructure serving millions of developers worldwide • Design and implement comprehensive monitoring, alerting, dashboards, metrics, and logging strategies • Architect and implement infrastructure automation using Terraform, Ansible, or Pulumi • Design and maintain CI/CD pipelines for reliable and consistent deployments • Create self-healing systems that respond automatically to common failure scenarios • Define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs) with product and engineering teams • Build systems to track and report reliability metrics • Lead incident response efforts and conduct thorough post-mortems • Develop and maintain runbooks for critical services • Build tools and processes to reduce Mean Time To Recovery (MTTR) • Identify and resolve infrastructure performance bottlenecks • Implement capacity planning strategies and optimize resource utilization • Reduce latency and improve system efficiency across global regions

🎯 Requirements

• 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering) • Strong programming skills in languages commonly used for automation (Python, Go, or similar) • Deep understanding of distributed systems • Experience with container orchestration platforms (Kubernetes) and cloud-native technologies • Proven track record of implementing and maintaining monitoring/observability solutions • Strong incident management skills with experience leading incident response • Experience with infrastructure as code and configuration management tools • Experience with Google Cloud Platform (GCP) services and tools (bonus) • Knowledge of modern observability platforms (Prometheus, Grafana, Datadog, etc.) (bonus) • Ability to approach complex operational challenges systematically and devise effective solutions • Capable of working independently while collaborating effectively with cross-functional teams • Strong communication skills to explain complex technical concepts to technical and non-technical audiences • Passion for staying current with industry best practices and new technologies • Strong belief in automating repetitive tasks and building self-healing systems • Legally authorized to work in the United States

🏖️ Benefits

• Competitive Salary & Equity • 401(k) Program with a 4% match (US Only) • Health, Dental, Vision and Life Insurance • Short Term and Long Term Disability • Paid Parental, Medical, Caregiver Leave • Flexible Time Off (FTO) + Holidays • Commuter Benefits (In-Office & US Only) • Monthly Wellness Stipend • Autonomous Work Environment • In Office Set-Up Reimbursement (In-Office Only) • Quarterly Team Gatherings • In Office Amenities (In-Office Only)

Apply Now

Similar Jobs

🔥 5 hours ago

TherapyNotes, LLC

51 - 200

💼 Consulting

⚖️ Legal

🏥 Healthcare

Site Reliability Engineer improving reliability, resilience, and observability for TherapyNotes’ behavioral health practice management and EHR SaaS platform. Automating operations across cloud infrastructure and 24×7 production services.

🔥 11 hours ago

Sphera

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Cybersecurity Engineer securing Sphera’s environmental, health, safety, and sustainability software for U.S. government and DoD clients. Managing RMF, STIG, ATO, vulnerability remediation, and DevSecOps compliance.

🔥 13 hours ago

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

FedRAMP SRE operating Arista Networks’ Kubernetes-native CloudVision networking SaaS. Ensuring reliable, secure, scalable production systems and leading infrastructure projects.

🔥 13 hours ago

Gainwell Technologies

10,000+ employees

💼 Consulting

📦 Logistics

⚕️ Healthcare Insurance

DevOps Engineer automating infrastructure, configuration management, and software releases. Supporting Gainwell’s cloud-based healthcare technology platforms and development teams.

🔥 13 hours ago

Gainwell Technologies

10,000+ employees

💼 Consulting

📦 Logistics

⚕️ Healthcare Insurance

DevOps Engineer automating configuration, releases, and cloud infrastructure for Gainwell Technologies’ healthcare platforms. Building scripts, CI processes, and infrastructure as code for SaaS and PaaS products.