Site Reliability Engineer

🔥 10 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of MyFitnessPal

MyFitnessPal

51 - 200 employees

Founded 2005

🏥 Healthcare

🍽️ Food & Beverage

💼 Consulting

💰 $18M Series A on 2013-08

Healthcare • Food & Beverage • Consulting

MyFitnessPal is a leading nutrition and fitness tracking application designed to help users reach their health and fitness goals. It offers an all-in-one solution for tracking food intake, exercise, and calories. With over 18 million foods in its database, MyFitnessPal allows users to track calories, macros, micronutrients, and more. The app integrates with many fitness devices to sync workouts, weight, and other health metrics. MyFitnessPal's personalized nutrition insights guide users towards building sustainable healthy habits. Available in both a free and premium version, it is praised for its user-friendly interface and effectiveness in helping users achieve weight loss and fitness objectives.

📋 Description

• Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue • Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform) • Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization • Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback) • Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review • Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time • Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk • Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function • Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable • Coach team members and engineers across the org on reliability patterns and operational best practices

🎯 Requirements

• 5+ years in site reliability, platform, or infrastructure engineering • Strong programming skills for automation and tooling (Go, Python, Typescript or similar) • Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus) • Proven track record leading incident response and building SLO-driven reliability practices • Working fluency with observability tooling (Datadog is a plus) • Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code • Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management) • The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win • Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus • Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus • Chaos engineering or game-day experience is a plus • Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus

🏖️ Benefits

• healthcare • parental planning • mental health benefits • annual performance bonus • a 401(k) plan and match • responsible time off • monthly wellness and technology allowances

Apply Now

Similar Jobs

🔥 21 hours ago

CXM

201 - 500

💸 Finance

💳 Fintech

Application Site Reliability Engineer focusing on .NET/C# services reliability for trading systems. Collaborating with software engineers to enhance service resilience and operational excellence.

AWS

Grafana

Postgres

Prometheus

Python

Terraform

.NET

🕒 Yesterday

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

DevOps Engineer supporting NVIDIA’s Rapids project for AI and data science initiatives. Collaborating with teams to ensure high-quality software releases and infrastructure maintenance.

AWS

Azure

Cloud

Docker

Jenkins

Linux

Open Source

Python

🕒 Yesterday

Softgic

51 - 200

💼 Consulting

🔒 Cybersecurity

DevOps Specialist managing Google Cloud infrastructure at Softgic. Automating and optimizing cloud systems and collaborating with development teams.

🗣️🇪🇸 Spanish Required

AWS

Cloud

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

Prometheus

Terraform

🕒 Yesterday

Global Enterprise Services, LLC (GES)

11 - 50

💼 Consulting

📦 Logistics

Reliability Engineer responsible for cloud platform performance and incident response, managing compliance. Requires strong technical expertise and 8 years of experience.

Cloud

🕒 Yesterday

IPolarity

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Site Reliability Engineer managing AWS GovCloud and multi-account environments. Building automation and overseeing cloud governance in a fully remote setup.

AWS

Cloud

Python

Ruby

Terraform

Go