Search Remote Jobs

Senior Site Reliability Engineer, Data & Analytics

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $102.8k - $190.2k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of XBOX

XBOX

10,000+ employees

Founded 2001

🎮 Gaming

🔧 Hardware

Gaming • Hardware

XBOX is Microsoft’s video gaming brand and platform encompassing consoles (XBOX Series X|S), first- and third-party games, subscription services (Xbox Game Pass), cloud gaming, an online store, accessories, and community/support features. The site highlights consoles, PC and mobile/handheld support, Game Pass day-one releases, hardware accessories, and developer resources, positioning Xbox as an integrated gaming hardware + software + services ecosystem and retail channel.

📋 Description

• Participate in an on-call rotation and drive incidents to resolution • Lead blameless postmortems and identify systemic reliability improvements • Partner with data, ML, and platform teams on batch, streaming, training, and inference workloads • Support ML training pipelines and inference services, including GPU workloads • Help define how data and ML services run on Kubernetes • Design and build automation and operational tooling, including workflows, diagnostic tooling, and runbooks • Build and evolve centralized platform services, shared tooling, data integrations, and access controls • Diagnose and resolve reliability, performance, and cost issues across distributed systems • Champion automation, documentation, and practices that reduce toil • Maintain infrastructure using Terraform and infrastructure-as-code principles • Improve CI/CD and GitOps workflows using Jenkins, GitHub Actions, and ArgoCD • Operate and improve containerized services on Kubernetes • Define and measure reliability using SLIs, SLOs, and error budgets • Run load tests, capacity modeling, and production validation • Build internal tools and paved paths that help teams operate safely and efficiently

🎯 Requirements

• Experience operating reliable, distributed systems in SRE, platform, or similar roles • Experience with data, analytics, ML, or large-scale distributed workloads • Strong knowledge of Linux, containers, Kubernetes, and cloud infrastructure • Experience building automation or internal tools using Python, Go, shell, or similar • Experience with infrastructure-as-code, such as Terraform • Experience with CI/CD or GitOps systems, such as Jenkins, GitHub Actions, or ArgoCD • Familiarity with observability, including metrics, logs, traces, alerting, and incident response • Solid understanding of SRE concepts, including SLIs, SLOs, error budgets, and postmortems • Experience using modern development and automation practices to improve reliability and efficiency • Experience building internal tooling, automation, or developer productivity systems • Strong communication skills with technical and cross-functional partners • Bonus: Experience with data and ML systems, including training pipelines, model serving, and GPU workloads • Bonus: Experience with distributed systems and messaging, such as Kafka or Pub/Sub • Bonus: Experience working in Kubernetes-based environments • Bonus: Familiarity with Prometheus and Grafana • Bonus: Experience operating systems in GCP or AWS cloud environments

🏖️ Benefits

• Medical, dental, and vision coverage • Health savings account or health reimbursement account • Healthcare spending accounts • Dependent care spending accounts • Life and AD&D insurance • Disability insurance • 401(k) with Company match • Tuition reimbursement • Charitable donation matching • Paid holidays and vacation • Paid sick time • Floating holidays • Compassion and bereavement leaves • Parental leave • Mental health and wellbeing programs • Fitness programs • Free and discounted games • Supplemental life and disability benefits • Legal service • ID protection • Rental insurance • Relocation assistance may be available if the Company requires geographic relocation • Incentive compensation may be available

Apply Now

Similar Jobs

🔥 35 minutes ago

Trumid

51 - 200

💳 Fintech

💸 Finance

☁️ SaaS

Senior DBRE safeguarding Postgres durability, recovery, and performance for Trumid’s fixed-income trading platform. Owning failover, observability, backups, and data resilience.

🔥 4 hours ago

Replicant

51 - 200

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Senior SRE building CI/CD, observability, and Kubernetes infrastructure for Replicant’s AI-powered customer-service platform. Improving reliability and developer tooling for large-scale conversational traffic.

🔥 4 hours ago

Ancestry

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

Site Reliability Engineer maintaining Ancestry’s family-history platform availability through incident response, AWS operations, and automation. Developing AI agentic workflows and tooling for a 24/7 Command Center.

🔥 6 hours ago

Red Alpha

51 - 200

🔒 Cybersecurity

🤖 Artificial Intelligence

Cloud DevOps Engineer automating Terraform, Python, Kubernetes, and GitOps deployments. Supporting Red Alpha’s secure cloud engineering solutions for a Federal Civilian-Health organization.

🔥 6 hours ago

Keyfactor

201 - 500

🚘 Automotive

🏥 Healthcare

🔐 Security

Senior DevOps Engineer automating cloud and on-premise infrastructure for Keyfactor’s cryptographic identity security platform. Improving Kubernetes, CI/CD, and Infrastructure as Code delivery.