Search Remote Jobs

Staff DevOps Engineer

Job not on LinkedIn

🔥 3 minutes ago

🇺🇸 United States – Remote

💵 $190k - $235k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cast & Crew

Cast & Crew

501 - 1000 employees

💼 Consulting

🏥 Healthcare

📦 Logistics

💰 Private equity on 2013-03

Consulting • Healthcare • Logistics

Cast & Crew is a provider of cloud-based software and professional services that support the full lifecycle of film, television, streaming and live-event productions. They offer production accounting, payroll and HR, timekeeping, digital onboarding, content and asset collaboration, reporting/analytics, and industry-specific services such as tax incentives guidance, workers' compensation, residuals, and financing. Cast & Crew serves production companies and entertainment businesses with B2B SaaS solutions and managed services to streamline workflows and ensure compliance across global productions.

📋 Description

• Architect and continuously improve Azure DevOps CI/CD pipelines, including pipeline-as-code standards, templating strategies, and artifact promotion workflows • Own the health and evolution of AWS EKS clusters, including node lifecycle, autoscaling, networking, RBAC, and upgrades • Design and enforce Infrastructure-as-Code practices and champion GitOps patterns • Drive platform reliability improvements using New Relic observability data in partnership with SRE • Define and maintain golden-path templates for containerized workloads, including Dockerfile standards and Helm chart libraries • Partner with engineering teams to onboard services and reduce toil through automation • Escalate and coordinate complex infrastructure incidents through PagerDuty, participate in on-call rotation, and lead post-incident reviews • Identify recurring failure modes and drive fixes that reduce page volume and MTTR • Maintain runbooks and platform documentation in Confluence • Define and socialize DevOps standards for pipelines, containers, secrets, and deployment safety • Conduct architecture reviews and provide technical guidance • Mentor senior and mid-level engineers through pairing, code review, and knowledge sharing • Identify tooling gaps and build business cases for platform investments

🎯 Requirements

• 8+ years of DevOps or platform engineering experience • At least 2 years operating at a Staff or Principal level in an organization of 100+ engineers • Deep, hands-on expertise with Kubernetes, specifically AWS EKS preferred, including workloads, networking, storage, and cluster operations at scale • Strong command of Azure DevOps Pipelines, including YAML pipeline authoring, library management, service connections, and environment promotion gates • Proven experience designing and maintaining CI/CD systems for microservice architectures with multiple independent teams • Experience operating observability platforms such as New Relic or Datadog to drive proactive reliability improvements • Proficiency in Python, Bash, or Go • Proficiency with Infrastructure-as-Code tooling such as Terraform, Pulumi, or CDK • Familiarity with feature flag patterns and progressive delivery; Unleash or equivalent is a plus • Excellent written communication skills • Experience with data engineering or ML infrastructure workloads on Kubernetes is preferred • Background contributing to or maintaining internal developer portals is preferred • Familiarity with FinOps practices and tooling for AWS cost attribution and optimization is preferred • Experience in SRE-adjacent roles and comfort with SLO/SLI definition and error budget policy is preferred

🏖️ Benefits

• Comprehensive medical, dental, and vision coverage • 401(k) match • Generous PTO • Paid parental leave • Health and wellness programs • Tuition reimbursement • Employee discounts

Apply Now

Similar Jobs

🔥 5 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Principal SRE architecting reliable network infrastructure for Akamai’s globally distributed cloud and edge platform. Automating operations, defining SLOs, and troubleshooting large-scale network systems.

🔥 23 hours ago

CDIT LLC

51 - 200

🤝 B2B

☁️ SaaS

Principal DevOps Architect designing secure AWS infrastructure for global multi-tenant SaaS services. Building Terraform, CI/CD, Kubernetes, observability, and HIPAA-compliant AI/ML platforms.

🕒 Yesterday

Camp Strategy

1 - 10

🏗️ Construction

📦 Logistics

🏨 Hospitality

Principal DevOps Engineer leading AWS, Kubernetes, and Terraform infrastructure for Campspot’s campground reservation software and camping marketplace. Driving automation, reliability, security, and cloud-cost optimization.

🕒 2 days ago

Fingerprint

51 - 200

🔒 Cybersecurity

🔌 API

☁️ SaaS

Staff SRE improving reliability for Fingerprint’s device-intelligence platform. Defining SLOs, strengthening incident operations, and enabling safe AI-assisted system operations.

🇺🇸 United States – Remote

💵 $177k - $240k / year

💰 $32M Series B on 2021-11

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 4 days ago

RealTime eClinical Solutions

51 - 200

🏥 Healthcare

🧬 Biotechnology

Principal DevOps Architect owning AWS, Terraform, CI/CD, observability, and AI/ML platforms. Ensuring HIPAA and SOC 2 compliance for clinical research SaaS.

🇺🇸 United States – Remote

💵 $155k - $195k / year

💰 Private Equity Round on 2022-01

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)