
10,000+ employees
Founded 2001
đŽ Gaming
đ§ Hardware
Gaming ⢠Hardware
XBOX is Microsoftâs video gaming brand and platform encompassing consoles (XBOX Series X|S), first- and third-party games, subscription services (Xbox Game Pass), cloud gaming, an online store, accessories, and community/support features. The site highlights consoles, PC and mobile/handheld support, Game Pass day-one releases, hardware accessories, and developer resources, positioning Xbox as an integrated gaming hardware + software + services ecosystem and retail channel.
đĽ 0 minutes ago
đşđ¸ United States â Remote
đľ $102.8k - $190.2k / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 0%
AWS
Cloud
Distributed Systems
Google Cloud Platform
Grafana
Jenkins
Kafka
Kubernetes
Linux
Prometheus
Python
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 2001
đŽ Gaming
đ§ Hardware
Gaming ⢠Hardware
XBOX is Microsoftâs video gaming brand and platform encompassing consoles (XBOX Series X|S), first- and third-party games, subscription services (Xbox Game Pass), cloud gaming, an online store, accessories, and community/support features. The site highlights consoles, PC and mobile/handheld support, Game Pass day-one releases, hardware accessories, and developer resources, positioning Xbox as an integrated gaming hardware + software + services ecosystem and retail channel.
⢠Participate in an on-call rotation and drive incidents to resolution ⢠Lead blameless postmortems and identify systemic reliability improvements ⢠Partner with data, ML, and platform teams on batch, streaming, training, and inference workloads ⢠Support ML training pipelines and inference services, including GPU workloads ⢠Help define how data and ML services run on Kubernetes ⢠Design and build automation and operational tooling, including workflows, diagnostic tooling, and runbooks ⢠Build and evolve centralized platform services, shared tooling, data integrations, and access controls ⢠Diagnose and resolve reliability, performance, and cost issues across distributed systems ⢠Champion automation, documentation, and practices that reduce toil ⢠Maintain infrastructure using Terraform and infrastructure-as-code principles ⢠Improve CI/CD and GitOps workflows using Jenkins, GitHub Actions, and ArgoCD ⢠Operate and improve containerized services on Kubernetes ⢠Define and measure reliability using SLIs, SLOs, and error budgets ⢠Run load tests, capacity modeling, and production validation ⢠Build internal tools and paved paths that help teams operate safely and efficiently
⢠Experience operating reliable, distributed systems in SRE, platform, or similar roles ⢠Experience with data, analytics, ML, or large-scale distributed workloads ⢠Strong knowledge of Linux, containers, Kubernetes, and cloud infrastructure ⢠Experience building automation or internal tools using Python, Go, shell, or similar ⢠Experience with infrastructure-as-code, such as Terraform ⢠Experience with CI/CD or GitOps systems, such as Jenkins, GitHub Actions, or ArgoCD ⢠Familiarity with observability, including metrics, logs, traces, alerting, and incident response ⢠Solid understanding of SRE concepts, including SLIs, SLOs, error budgets, and postmortems ⢠Experience using modern development and automation practices to improve reliability and efficiency ⢠Experience building internal tooling, automation, or developer productivity systems ⢠Strong communication skills with technical and cross-functional partners ⢠Bonus: Experience with data and ML systems, including training pipelines, model serving, and GPU workloads ⢠Bonus: Experience with distributed systems and messaging, such as Kafka or Pub/Sub ⢠Bonus: Experience working in Kubernetes-based environments ⢠Bonus: Familiarity with Prometheus and Grafana ⢠Bonus: Experience operating systems in GCP or AWS cloud environments
⢠Medical, dental, and vision coverage ⢠Health savings account or health reimbursement account ⢠Healthcare spending accounts ⢠Dependent care spending accounts ⢠Life and AD&D insurance ⢠Disability insurance ⢠401(k) with Company match ⢠Tuition reimbursement ⢠Charitable donation matching ⢠Paid holidays and vacation ⢠Paid sick time ⢠Floating holidays ⢠Compassion and bereavement leaves ⢠Parental leave ⢠Mental health and wellbeing programs ⢠Fitness programs ⢠Free and discounted games ⢠Supplemental life and disability benefits ⢠Legal service ⢠ID protection ⢠Rental insurance ⢠Relocation assistance may be available if the Company requires geographic relocation ⢠Incentive compensation may be available
Apply NowđĽ 35 minutes ago
Senior DBRE safeguarding Postgres durability, recovery, and performance for Trumidâs fixed-income trading platform. Owning failover, observability, backups, and data resilience.
đĽ 4 hours ago
Senior SRE building CI/CD, observability, and Kubernetes infrastructure for Replicantâs AI-powered customer-service platform. Improving reliability and developer tooling for large-scale conversational traffic.
đĽ 4 hours ago
Site Reliability Engineer maintaining Ancestryâs family-history platform availability through incident response, AWS operations, and automation. Developing AI agentic workflows and tooling for a 24/7 Command Center.
đşđ¸ United States â Remote
đľ $106k - $118.5k / year
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đĽ 6 hours ago
Cloud DevOps Engineer automating Terraform, Python, Kubernetes, and GitOps deployments. Supporting Red Alphaâs secure cloud engineering solutions for a Federal Civilian-Health organization.
đşđ¸ United States â Remote
đľ $180k - $240k / year
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đĽ 6 hours ago
Senior DevOps Engineer automating cloud and on-premise infrastructure for Keyfactorâs cryptographic identity security platform. Improving Kubernetes, CI/CD, and Infrastructure as Code delivery.