Principal Site Reliability Engineer

🕒 Yesterday

🇺🇸 United States – Remote

💵 $165k - $185k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tandem Diabetes Care

Tandem Diabetes Care

1001 - 5000 employees

🏥 Healthcare

🏭 Manufacturing

🔧 Hardware

Healthcare • Manufacturing • Hardware

Tandem Diabetes Care is a medical device company that designs and manufactures insulin pumps and automated insulin delivery systems powered by its Control-IQ+ predictive algorithm. Its products, including the t:slim X2 and Tandem Mobi, integrate with continuous glucose monitors and compatible smartphones, support remote software updates, and are accompanied by mobile and cloud-based applications, training, and 24/7 support to help people manage insulin therapy.

📋 Description

• Lead day-to-day production support, including intake, triage, prioritization, escalation, queue health, and change execution • Establish consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and ownership of open issues • Lead incident management end-to-end, including incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure • Coordinate production security incident response with Security teams • Own on-call strategy, rotation design, escalation paths, alert tuning, and tooling such as PagerDuty and New Relic • Serve as a senior escalation tier for high-severity incidents • Build runbooks for common failure modes and first-line resolution • Define and own SLIs and SLOs for critical services • Reduce MTTD and MTTR through instrumentation, alerting, diagnostics, and automation • Convert recurring support burden into permanent fixes, automation, or documentation • Maintain technology currency and lifecycle management for cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components • Own business continuity and disaster recovery readiness, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and RTO/RPO objectives • Lead infrastructure automation with Terraform • Eliminate toil through automation • Add reliability guardrails to CI/CD pipelines, including automated rollback, change-risk checks, and progressive delivery • Maintain production systems in accordance with regulatory and compliance requirements • Maintain business continuity and disaster recovery documentation and audit evidence • Partner with Security, Quality, and Compliance teams on audits, compliance, and remediation • Grow SRE and DevOps engineers through pairing, design and code review, and incident debriefs • Foster open communication among junior and contract engineers • Communicate documentation-first across distributed, multi-time-zone teams • Partner with software engineering, QA, and architecture to embed reliability into the development lifecycle • Inform capacity planning and scaling strategy with the Test team • Introduce proactive resilience testing such as game days • Align business continuity and disaster recovery capabilities with application requirements • Support cloud cost optimization through rightsizing, reserved capacity, and observability spend governance • Ensure work complies with company policies and applicable Privacy/HIPAA, regulatory, legal, and safety requirements

🎯 Requirements

• Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events • Strong grounding in SRE principles, including SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering • Demonstrated experience owning on-call strategy, including rotation design, alert tuning, and escalation • Expertise with Terraform or comparable IaC at scale, including module design, state management, and policy-as-code guardrails • Hands-on experience building CI/CD pipelines with reliability guardrails using GitHub Actions, Octopus Deploy, or Azure DevOps • Deep experience with at least one major cloud platform: AWS, Azure, or GCP • Experience with Docker and Kubernetes • Working knowledge of observability tooling such as Prometheus, Grafana, Datadog, CloudWatch, and ELK/OpenSearch • Experience designing and testing disaster recovery, including backup/restore, failover, and RTO/RPO validation • Working knowledge of cloud security and compliance practices, including IAM, network segmentation, encryption, vulnerability and patch management, and cloud cost optimization • Proficiency in at least one scripting or programming language such as Python, Go, or Bash • Experience in FDA and ISO regulated industries and agile methodologies preferred • B.S. in Computer Science or equivalent combination of education and applicable job experience, including technical school training and certifications in networks, servers, and cloud infrastructure; demonstrated production experience weighs more heavily than degree • Relevant cloud certifications such as AWS/Azure/GCP Professional or Architect level preferred • 10+ years in Site Reliability Engineering, DevOps, or infrastructure engineering • 2+ years mentoring or technically leading other engineers, including remote, offshore, or contracted partner engineers • Must be within the United States • Successful completion of pre-employment drug test and background check • Compliance with applicable company, Privacy/HIPAA, regulatory, legal, and safety requirements

🏖️ Benefits

• Medical, dental, and vision benefits available the first day • Health savings accounts • Flexible savings accounts • 11 paid holidays per year • Minimum of 20 days of paid time off, with accrual starting on day 1 • 401(k) plan with company match • Employee Stock Purchase plan • Equipment provided • Virtual training • Bonus and competitive compensation package • Joy-focused workplace supporting well-being, achievement, growth, fun, and camaraderie

Apply Now

Similar Jobs

🕒 Yesterday

Bixal

51 - 200

📣 Marketing

📦 Logistics

🏥 Healthcare

Director of DevSecOps leading secure cloud platforms and DevSecOps practice growth for Bixal’s federal government consulting clients. Recruiting teams, guiding architecture, and supporting compliant delivery and proposals.

🕒 Yesterday

PhoenixTeam

51 - 200

💳 Fintech

🏠 Real Estate

🤖 Artificial Intelligence

DevOps Manager modernizing Jenkins-based CI/CD and Fortify quality controls for PhoenixTeam's federal FHA mortgage program. Coordinating releases, documentation, and delivery across development teams.

🕒 2 days ago

Cast & Crew

501 - 1000

☁️ SaaS

📱 Media

👥 HR Tech

Staff DevOps Engineer owning AWS EKS, Azure DevOps CI/CD, and platform reliability. Supporting Cast & Crew’s entertainment technology and services business through automation and technical leadership.

🇺🇸 United States – Remote

💵 $190k - $235k / year

💰 Private equity on 2013-03

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

Cast & Crew

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Staff DevOps Engineer owning AWS EKS, Azure DevOps CI/CD, and developer tooling. Improving platform reliability for Cast & Crew’s entertainment technology and services business.

🕒 2 days ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Principal SRE architecting reliable network infrastructure for Akamai’s globally distributed cloud and edge platform. Automating operations, defining SLOs, and troubleshooting large-scale network systems.