Site Reliability Engineer

🕒 July 26

🇺🇸 United States – Remote

💵 $120k - $165k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 9%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of MyFitnessPal

MyFitnessPal

51 - 200 employees

Founded 2005

🏥 Healthcare

🍽️ Food & Beverage

💼 Consulting

💰 $18M Series A on 2013-08

Healthcare • Food & Beverage • Consulting

MyFitnessPal is a leading nutrition and fitness tracking application designed to help users reach their health and fitness goals. It offers an all-in-one solution for tracking food intake, exercise, and calories. With over 18 million foods in its database, MyFitnessPal allows users to track calories, macros, micronutrients, and more. The app integrates with many fitness devices to sync workouts, weight, and other health metrics. MyFitnessPal's personalized nutrition insights guide users towards building sustainable healthy habits. Available in both a free and premium version, it is praised for its user-friendly interface and effectiveness in helping users achieve weight loss and fitness objectives.

📋 Description

• Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue • Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform) • Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization • Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback) • Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review • Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time • Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk • Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function • Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable • Coach team members and engineers across the org on reliability patterns and operational best practices

🎯 Requirements

• 5+ years in site reliability, platform, or infrastructure engineering • Strong programming skills for automation and tooling (Go, Python, Typescript or similar) • Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus) • Proven track record leading incident response and building SLO-driven reliability practices • Working fluency with observability tooling (Datadog is a plus) • Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code • Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management) • The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win • Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus • Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus • Chaos engineering or game-day experience is a plus • Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus

🏖️ Benefits

• healthcare • parental planning • mental health benefits • annual performance bonus • a 401(k) plan and match • responsible time off • monthly wellness and technology allowances

Apply Now

Similar Jobs

🕒 July 25

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer handling high-performance storage for AI platforms at Mirantis. Integrating and operating storage solutions within Kubernetes environments.

🕒 July 25

Chromatic (we're hiring!)

11 - 50

☁️ SaaS

⚡ Productivity

🏢 Enterprise

DevOps Engineer focused on securing and stabilizing Chromatic's AWS infrastructure. Collaborating across teams to enhance security posture and developer workflows within a remote environment.

🕒 July 24

Compassion International

1001 - 5000

🤝 Non-profit

🤲 Charity

🌍 Social Impact

Sr. DevOps Engineer managing scalable AWS platforms, automation, Kubernetes, and AI/ML infrastructure. Supporting Compassion International’s mission to release children from poverty in Jesus’ name.

🕒 July 24

Abacus Insights

51 - 200

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior Engineering Manager leading DevOps, infrastructure, and release engineering for Abacus Insights’ healthcare data SaaS platform. Improving CI/CD reliability, cloud scalability, observability, and safe delivery.

🕒 July 24

Juniper Square

201 - 500

💸 Finance

🏠 Real Estate

☁️ SaaS

Senior SRE scaling reliability and observability for Juniper Square’s private-markets technology platform. Owning Kubernetes, AWS, CI/CD, infrastructure automation, and cross-team reliability standards.