
51 - 200 employees
Founded 2022
π€ Artificial Intelligence
βοΈ SaaS
π€ B2B
π° $20M Seed on 2024-06
Artificial Intelligence β’ SaaS β’ B2B
Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.
π₯ 1 hour ago
π California β Remote
π΅ $150k - $200k / year
β° Full Time
π‘ Mid-level
π Senior
β DevOps & Site Reliability Engineer (SRE)
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
Founded 2022
π€ Artificial Intelligence
βοΈ SaaS
π€ B2B
π° $20M Seed on 2024-06
Artificial Intelligence β’ SaaS β’ B2B
Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.
β’ Define and implement SLIs/SLOs for critical services β’ Lead incident response and coordinate cross-team mitigation efforts β’ Conduct blameless postmortems and ensure corrective actions are completed β’ Perform production readiness reviews for new services and features β’ Identify systemic risks and drive preventative improvements β’ Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) β’ Build internal tooling for reliability tracking and reporting β’ Automate recurring operational workflows β’ Strengthen CI/CD reliability and release processes β’ Partner with engineering teams to improve system resilience β’ Provide guidance on fault tolerance, scalability, and failure handling. β’ Contribute to architectural discussions with a reliability-first mindset.
β’ 5+ years of experience in SRE, Reliability Engineering, or Production Engineering β’ Strong Linux systems and Networking expertise β’ Experience managing containerized production systems β’ Strong understanding of distributed systems and failure modes β’ Experience defining and managing SLIs/SLOs β’ Proven incident response and postmortem leadership experience β’ Strong scripting or programming skills β’ Experience with monitoring and alerting systems β’ Excellent written communication skills β’ Successful completion of a background check. β’ Preferred: Experience with GPU infrastructure or AI/ML platforms β’ Experience improving reliability in high-growth or large scale environments β’ Familiarity with GPU observability tooling β’ Experience with Infrastructure as Code β’ Experience working in startup environments β’ Experience building internal reliability platforms or frameworks.
β’ Meaningful equity in a fast-growing company- everyone on the team receives stock options β your impact drives our growth, and you share in the upside. β’ Generous medical, dental & vision plans β’ Flexible PTO- take the time you need to recharge β’ Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication β’ Join a passionate team on the cutting edge of AI infrastructure β where culture, learning, and ownership are at the heart of how we scale.
Apply Nowπ₯ 1 hour ago
Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.
πΊπΈ United States β Remote
π΅ $160k - $185k / year
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)
π₯ 1 hour ago
DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.
π₯ 2 hours ago
Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.
πΊπΈ United States β Remote
π΅ $169k - $215k / year
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)
π₯ 2 hours ago
Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.
πΊπΈ United States β Remote
π΅ $179.4k - $232.1k / year
π° $75M Debt Financing - Thumbtack on 2024-07
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)
π₯ 2 hours ago
Senior Principal DevSecOps Engineer designing and implementing DevSecOps platforms for Collins Aerospace. Collaborating with engineers and cybersecurity professionals to enhance software development pipelines and processes.
πΊπΈ United States β Remote
π΅ $132.4k - $251.6k / year
π° $200k Grant - RTX on 2024-11
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)