Search Remote Jobs

Site Reliability Engineer

πŸ”₯ 1 hour ago

Apply Now
Find Similar Remote Jobs

πŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

πŸ€– Artificial Intelligence

☁️ SaaS

🀝 B2B

πŸ’° $20M Seed on 2024-06

Artificial Intelligence β€’ SaaS β€’ B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

πŸ“‹ Description

β€’ Define and implement SLIs/SLOs for critical services β€’ Lead incident response and coordinate cross-team mitigation efforts β€’ Conduct blameless postmortems and ensure corrective actions are completed β€’ Perform production readiness reviews for new services and features β€’ Identify systemic risks and drive preventative improvements β€’ Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) β€’ Build internal tooling for reliability tracking and reporting β€’ Automate recurring operational workflows β€’ Strengthen CI/CD reliability and release processes β€’ Partner with engineering teams to improve system resilience β€’ Provide guidance on fault tolerance, scalability, and failure handling. β€’ Contribute to architectural discussions with a reliability-first mindset.

🎯 Requirements

β€’ 5+ years of experience in SRE, Reliability Engineering, or Production Engineering β€’ Strong Linux systems and Networking expertise β€’ Experience managing containerized production systems β€’ Strong understanding of distributed systems and failure modes β€’ Experience defining and managing SLIs/SLOs β€’ Proven incident response and postmortem leadership experience β€’ Strong scripting or programming skills β€’ Experience with monitoring and alerting systems β€’ Excellent written communication skills β€’ Successful completion of a background check. β€’ Preferred: Experience with GPU infrastructure or AI/ML platforms β€’ Experience improving reliability in high-growth or large scale environments β€’ Familiarity with GPU observability tooling β€’ Experience with Infrastructure as Code β€’ Experience working in startup environments β€’ Experience building internal reliability platforms or frameworks.

πŸ–οΈ Benefits

β€’ Meaningful equity in a fast-growing company- everyone on the team receives stock options β€” your impact drives our growth, and you share in the upside. β€’ Generous medical, dental & vision plans β€’ Flexible PTO- take the time you need to recharge β€’ Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication β€’ Join a passionate team on the cutting edge of AI infrastructure β€” where culture, learning, and ownership are at the heart of how we scale.

Apply Now

Similar Jobs

πŸ”₯ 1 hour ago

Talkiatry

501 - 1000

πŸ₯ Healthcare

πŸ‘₯ B2C

Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.

πŸ”₯ 1 hour ago

Careerswift

2 - 10

πŸ‘₯ HR Tech

🎯 Recruiter

☁️ SaaS

DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.

πŸ”₯ 2 hours ago

Multi Media, LLC

51 - 200

πŸ’Ό Consulting

πŸ“£ Marketing

πŸ“± Media

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

πŸ”₯ 2 hours ago

Thumbtack

1001 - 5000

πŸͺ Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

πŸ‡ΊπŸ‡Έ United States – Remote

πŸ’΅ $179.4k - $232.1k / year

πŸ’° $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

β›‘ DevOps & Site Reliability Engineer (SRE)

πŸ”₯ 2 hours ago

RTX

10,000+ employees

πŸš€ Aerospace

πŸŽ–οΈ Defense

🏭 Manufacturing

Senior Principal DevSecOps Engineer designing and implementing DevSecOps platforms for Collins Aerospace. Collaborating with engineers and cybersecurity professionals to enhance software development pipelines and processes.