Staff Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🗽 New York – Remote

infoinfo

💵 $241k - $270k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 15%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Garner Health

Garner Health

51 - 200 employees

💼 Consulting

📦 Logistics

🏥 Healthcare

Consulting • Logistics • Healthcare

Garner Health is focused on improving the way employees find high-quality doctors. With a belief in transparency and data-driven decision-making, they create solutions that facilitate better medical choices for employees. The team comprises healthcare operators, clinicians, engineers, and benefits experts, allowing for a multidisciplinary approach in developing these healthcare solutions.

📋 Description

• Own the end-to-end reliability, performance, and resilience strategy for Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads • Architect the SLO framework for critical services and lead technical decision-making for scale • Lead and improve the incident response program, including on-call participation, complex escalations, root cause analysis, and corrective actions • Architect monitoring, alerting, and observability platforms • Transform ambiguous scaling and reliability requirements into automated, composable Terraform infrastructure-as-code deliverables • Identify and implement cloud cost-efficiency and performance improvements • Redirect inefficient or fragile infrastructure workflows and reduce impactful technical debt • Use AI tools and automation to convert repetitive operational work into hands-free, monitored processes • Build deployment and observability standards that help engineers ship AI features reliably • Mentor engineers and provide feedback on operational rigor • Ensure infrastructure and operations meet security and HIPAA compliance obligations • Lead rigorous reviews of infrastructure changes

🎯 Requirements

• 7+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role • Deep expertise with Kubernetes and Terraform in a cloud-first environment; AWS preferred • Experience designing SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews • Strong Python or Go skills applied to infrastructure automation • Kubernetes API experience is a plus • Track record driving cloud cost-efficiency and performance optimization across compute, storage, and networking • Mentorship experience and ability to set technical direction as the senior reliability voice • Excellent communication skills with technical and non-technical stakeholders • Fluency with AI tools such as Claude applied to engineering and operations workflows, or strong motivation to build it fast • Experience supporting AI/ML or data-intensive workloads in production is a plus • Experience in a security-conscious or regulated environment such as HIPAA or SOC 2 is a plus • Technologies used include AWS, Kubernetes, Terraform, Istio, Python, Go, TypeScript, Postgres, NATS, Datadog, and GitLab

🏖️ Benefits

• Equity incentive plan • Flexible PTO • Medical plan options • Dental plan options • Vision plan options • 401(k) • Teladoc Health • Remote work with occasional travel to HQ

Apply Now

Similar Jobs

🕒 2 days ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

DevSecOps Security Engineer automating AWS cloud security, compliance, and vulnerability workflows for GDIT’s U.S. government missions. Managing continuous ATO pipelines, IaC security, and audit evidence at scale.

🕒 4 days ago

TEKsystems

10,000+ employees

💼 Consulting

🎯 Recruiter

🏢 Enterprise

CloudOps Practice Architect designing AWS/GCP cloud-native platforms and SRE solutions for TEKsystems Global Services. Leading hands-on architecture, automation, reliability, and global team mentorship.

🕒 4 days ago

TEKsystems

10,000+ employees

💼 Consulting

🎯 Recruiter

🏢 Enterprise

CloudOps Practice Architect designing AWS/GCP SRE platforms and agentic architectures for TEKsystems enterprise clients. Mentoring global teams and accelerating cloud solution delivery from concept to production.

🕒 4 days ago

Identity Digital Inc.

201 - 500

💼 Consulting

📣 Marketing

🛍️ eCommerce

Staff DevOps Engineer owning cloud-native infrastructure for Identity Digital’s DNS-native AI-agent identity platform. Building enterprise production readiness, DNS operations, observability, security, and on-call practices.

🕒 4 days ago

Identity Digital Inc.

201 - 500

💼 Consulting

📣 Marketing

🛍️ eCommerce

Staff DevOps Engineer owning cloud infrastructure, DNS, reliability, and delivery pipelines for Identity Digital’s AI-agent identity platform. Building enterprise-ready operations, security, observability, and on-call practices.