
51 - 200 employees
Founded 2022
🤖 Artificial Intelligence
☁️ SaaS
🤝 B2B
💰 $20M Seed on 2024-06
Artificial Intelligence • SaaS • B2B
Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.
🕒 4 days ago
🏄 California – Remote
💵 $150k - $200k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
Founded 2022
🤖 Artificial Intelligence
☁️ SaaS
🤝 B2B
💰 $20M Seed on 2024-06
Artificial Intelligence • SaaS • B2B
Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.
• Define and implement SLIs/SLOs for critical services • Lead incident response and coordinate cross-team mitigation efforts • Conduct blameless postmortems and ensure corrective actions are completed • Perform production readiness reviews for new services and features • Identify systemic risks and drive preventative improvements • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) • Build internal tooling for reliability tracking and reporting • Automate recurring operational workflows • Strengthen CI/CD reliability and release processes • Partner with engineering teams to improve system resilience • Provide guidance on fault tolerance, scalability, and failure handling. • Contribute to architectural discussions with a reliability-first mindset.
• 5+ years of experience in SRE, Reliability Engineering, or Production Engineering • Strong Linux systems and Networking expertise • Experience managing containerized production systems • Strong understanding of distributed systems and failure modes • Experience defining and managing SLIs/SLOs • Proven incident response and postmortem leadership experience • Strong scripting or programming skills • Experience with monitoring and alerting systems • Excellent written communication skills • Successful completion of a background check. • Preferred: Experience with GPU infrastructure or AI/ML platforms • Experience improving reliability in high-growth or large scale environments • Familiarity with GPU observability tooling • Experience with Infrastructure as Code • Experience working in startup environments • Experience building internal reliability platforms or frameworks.
• Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. • Generous medical, dental & vision plans • Flexible PTO- take the time you need to recharge • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.
Apply Now🕒 4 days ago
Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.
🇺🇸 United States – Remote
💵 $160k - $185k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 4 days ago
Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.
🇺🇸 United States – Remote
💵 $169k - $215k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 5 days ago
Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.
🇺🇸 United States – Remote
💵 $179.4k - $232.1k / year
💰 $75M Debt Financing - Thumbtack on 2024-07
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 5 days ago
Lead Power Platform Reliability Engineer enhancing enterprise-level solutions through collaboration and mentorship. Shape future data-driven applications and drive cloud integration.
🕒 5 days ago
DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.
🇺🇸 United States – Remote
💵 $108.5k - $184.4k / year
💰 $29M Grant on 2023-03
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)