Principal Site Reliability Engineer – Temp to Hire

🔥 13 hours ago

🇺🇸 United States – Remote

💵 $165k - $185k / year

⏳ Contract/Temporary

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tandem Diabetes Care

Tandem Diabetes Care

1001 - 5000 employees

🏥 Healthcare

🏭 Manufacturing

🔧 Hardware

Healthcare • Manufacturing • Hardware

Tandem Diabetes Care is a medical device company that designs and manufactures insulin pumps and automated insulin delivery systems powered by its Control-IQ+ predictive algorithm. Its products, including the t:slim X2 and Tandem Mobi, integrate with continuous glucose monitors and compatible smartphones, support remote software updates, and are accompanied by mobile and cloud-based applications, training, and 24/7 support to help people manage insulin therapy.

📋 Description

• Lead day-to-day production support, including intake, triage, prioritization, escalation, queue health, and change execution • Establish consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and ownership of open issues • Lead incident management end-to-end, including incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure • Coordinate response to production security incidents with Security teams • Own on-call strategy, rotation design, escalation paths, alert tuning, and tooling such as PagerDuty and New Relic • Build runbooks that standardize responses and enable first-line resolution • Define and own SLIs and SLOs for critical services • Reduce MTTD and MTTR through instrumentation, alerting, diagnostics, and automation • Convert recurring support burden into permanent fixes, automation, or documentation • Maintain technology currency and lifecycle management for production platforms, including cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components • Own business continuity and disaster recovery readiness, including backup and recovery strategies, recovery testing, failover capabilities, runbooks, and RTO/RPO objectives • Lead infrastructure automation with Terraform • Eliminate toil and reduce manual operational work through automation • Add reliability guardrails to CI/CD pipelines, including automated rollback, change-risk checks, and progressive delivery • Maintain production systems according to regulatory and compliance requirements and ensure audit readiness • Maintain business continuity and disaster recovery documentation and evidence • Partner with Security, Quality, and Compliance teams on regulatory compliance, audits, and remediation • Grow SRE and DevOps engineers through pairing, design and code review, and incident debriefs • Promote documentation-first communication across distributed, multi-time-zone teams • Partner with software engineering, QA, and architecture to embed reliability and operability into development • Inform capacity planning and scaling strategy with the Test team • Introduce proactive resilience testing such as game days • Align business continuity and disaster recovery capabilities with application requirements and recovery objectives • Support cloud cost optimization, including rightsizing, reserved capacity, and observability spend governance • Ensure work complies with Privacy/HIPAA and other regulatory, legal, and safety requirements

🎯 Requirements

• Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events • Strong grounding in SRE principles: SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering • Experience owning on-call strategy, including rotation design, alert tuning, and escalation • Expertise with Terraform or comparable IaC at scale, including module design, state management, and policy-as-code guardrails • Hands-on experience building CI/CD pipelines with reliability guardrails using GitHub Actions, Octopus Deploy, or Azure DevOps • Deep experience with at least one major cloud platform: AWS, Azure, or GCP • Experience with Docker and Kubernetes • Working knowledge of observability tooling such as Prometheus, Grafana, Datadog, CloudWatch, or ELK/OpenSearch • Experience designing and testing disaster recovery, including backup/restore, failover, and RTO/RPO validation • Working knowledge of cloud security and compliance practices, including IAM, network segmentation, encryption, vulnerability and patch management, and cloud cost optimization • Proficiency in at least one scripting or programming language, such as Python, Go, or Bash • 10+ years in Site Reliability Engineering, DevOps, or infrastructure engineering • 2+ years mentoring or technically leading other engineers, including remote, offshore, or contracted engineers • Experience in FDA and ISO regulated industries and agile methodologies preferred • B.S. in Computer Science or equivalent combination of education and applicable job experience, including technical school training and certifications; demonstrated production experience weighs more heavily than degree • Relevant cloud certifications preferred • Must be authorized to work for any employer in the U.S.; employer cannot sponsor or take over sponsorship of an employment visa

🏖️ Benefits

• Equipment for the role will be provided • Training will occur virtually • Competitive compensation package including bonus • Robust benefits package • Tandem-sponsored benefits contingent upon conversion from temporary to regular full-time status • Equal opportunity and inclusive workplace • Joyful workplace celebrating achievements and supporting well-being

Apply Now

Similar Jobs

🕒 August 19

3Core Systems, Inc

51 - 200

🤝 B2B

💼 Consulting

👥 HR Tech

DevOps Engineer managing Redwood RunMyJobs scheduling and automation for the Texas Railroad Commission. Optimizing workflows, integrations, reliability, and SLA adherence remotely.

🕒 August 18

Staff SRE/Cloud SME leading ASCENDING’s transformation of monolithic systems into resilient, cloud-native AWS and Azure architectures. Designing Kubernetes infrastructure, cloud networking, monitoring, security, and DevOps practices while mentoring engineers.