Site Reliability Engineer

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

💰 $20M Seed on 2024-06

Artificial Intelligence • SaaS • B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

📋 Description

• Define and implement SLIs/SLOs for critical services • Lead incident response and coordinate cross-team mitigation efforts • Conduct blameless postmortems and ensure corrective actions are completed • Perform production readiness reviews for new services and features • Identify systemic risks and drive preventative improvements • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) • Build internal tooling for reliability tracking and reporting • Automate recurring operational workflows • Strengthen CI/CD reliability and release processes • Partner with engineering teams to improve system resilience • Provide guidance on fault tolerance, scalability, and failure handling. • Contribute to architectural discussions with a reliability-first mindset.

🎯 Requirements

• 5+ years of experience in SRE, Reliability Engineering, or Production Engineering • Strong Linux systems and Networking expertise • Experience managing containerized production systems • Strong understanding of distributed systems and failure modes • Experience defining and managing SLIs/SLOs • Proven incident response and postmortem leadership experience • Strong scripting or programming skills • Experience with monitoring and alerting systems • Excellent written communication skills • Successful completion of a background check. • Preferred: Experience with GPU infrastructure or AI/ML platforms • Experience improving reliability in high-growth or large scale environments • Familiarity with GPU observability tooling • Experience with Infrastructure as Code • Experience working in startup environments • Experience building internal reliability platforms or frameworks.

🏖️ Benefits

• Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. • Generous medical, dental & vision plans • Flexible PTO- take the time you need to recharge • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

Apply Now

Similar Jobs

🔥 1 hour ago

Talkiatry

501 - 1000

🏥 Healthcare

👥 B2C

Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.

AWS

Grafana

Prometheus

Python

Terraform

TypeScript

🔥 1 hour ago

Careerswift

2 - 10

👥 HR Tech

🎯 Recruiter

☁️ SaaS

DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.

AWS

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

🔥 2 hours ago

Multi Media, LLC

51 - 200

💼 Consulting

📣 Marketing

📱 Media

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

Ansible

Cloud

Django

Docker

Flask

Java

Kubernetes

Linux

Laravel

Python

Rust

Switching

Terraform

Go

🔥 3 hours ago

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 United States – Remote

💵 $179.4k - $232.1k / year

💰 $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

Distributed Systems

DNS

JavaScript

Linux

Microservices

PHP

Python

SDLC

TCP/IP

Go

🔥 3 hours ago

RTX

10,000+ employees

🚀 Aerospace

🎖️ Defense

🏭 Manufacturing

Senior Principal DevSecOps Engineer designing and implementing DevSecOps platforms for Collins Aerospace. Collaborating with engineers and cybersecurity professionals to enhance software development pipelines and processes.

Docker

Jenkins

Kubernetes

Linux

Maven

Perl

Python

VMware