Search Remote Jobs

Staff DevOps Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $190k - $235k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cast & Crew

Cast & Crew

501 - 1000 employees

☁️ SaaS

📱 Media

👥 HR Tech

💰 Private equity on 2013-03

SaaS • Media • HR Tech

Cast & Crew is a provider of cloud-based software and services that support the full production lifecycle for the entertainment industry. The company offers production accounting and AP software (PSL+), digital onboarding and timekeeping tools (Start+, Hours+), collaboration and content tools (Studio+), reporting/data integrations, and a full-service payroll and financial services suite including residuals, workers' compensation, tax-incentive guidance, and related production HR resources. Cast & Crew serves film, television, streaming and live entertainment productions, delivering B2B SaaS products plus specialized payroll and workforce support for production employers and crews.

📋 Description

• Architect and continuously improve Azure DevOps CI/CD pipelines, including pipeline-as-code standards, templating strategies, and artifact promotion workflows • Own the health and evolution of AWS EKS clusters, including node lifecycle, autoscaling, networking, RBAC, and cluster upgrades • Design and enforce Infrastructure-as-Code practices and champion GitOps patterns • Drive platform reliability improvements using observability data from New Relic in partnership with SRE • Define and maintain golden-path templates for containerized workloads, including Dockerfile standards and Helm chart libraries • Partner with engineering teams to onboard new services and reduce toil through automation • Serve as an escalation point for complex infrastructure incidents through PagerDuty and participate in the on-call rotation • Lead post-incident reviews and drive systemic fixes that reduce page volume and MTTR • Maintain runbooks and platform documentation in Confluence • Define and socialize DevOps standards across the engineering organization • Conduct architecture reviews and provide technical guidance on infrastructure decisions • Mentor senior and mid-level engineers through pairing, code review, and knowledge sharing • Identify tooling gaps and build business cases for platform investments

🎯 Requirements

• 8+ years of DevOps or platform engineering experience • At least 2 years operating at a Staff or Principal level in an organization of 100+ engineers • Deep, hands-on expertise with Kubernetes; EKS specifically preferred • Experience troubleshooting workloads, networking, storage, and cluster operations at scale • Strong command of Azure DevOps Pipelines, including YAML pipeline authoring, library management, service connections, and environment promotion gates • Proven experience designing and maintaining CI/CD systems for microservice architectures with multiple independent teams • Experience operating observability platforms such as New Relic or Datadog to drive proactive reliability improvements • Proficiency in at least one scripting language: Python, Bash, or Go • Proficiency with Infrastructure-as-Code tooling: Terraform, Pulumi, or CDK • Familiarity with feature flag patterns and progressive delivery; Unleash or equivalent is a plus • Excellent written communication skills and ability to translate complex infrastructure decisions into practical guidance • Preferred: experience with data engineering or ML infrastructure workloads on Kubernetes • Preferred: background contributing to or maintaining internal developer portals such as Backstage • Preferred: familiarity with FinOps practices and AWS cost attribution and optimization • Preferred: experience in SRE-adjacent roles and comfort with SLO/SLI definition and error budget policy

🏖️ Benefits

• Comprehensive medical coverage • Dental coverage • Vision coverage • 401(k) match • Generous PTO • Paid parental leave • Health and wellness programs • Tuition reimbursement • Employee discounts

Apply Now

Similar Jobs

🔥 3 minutes ago

Cast & Crew

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Staff DevOps Engineer owning AWS EKS, Azure DevOps CI/CD, and developer tooling. Improving platform reliability for Cast & Crew’s entertainment technology and services business.

🔥 5 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Principal SRE architecting reliable network infrastructure for Akamai’s globally distributed cloud and edge platform. Automating operations, defining SLOs, and troubleshooting large-scale network systems.

🔥 23 hours ago

CDIT LLC

51 - 200

🤝 B2B

☁️ SaaS

Principal DevOps Architect designing secure AWS infrastructure for global multi-tenant SaaS services. Building Terraform, CI/CD, Kubernetes, observability, and HIPAA-compliant AI/ML platforms.

🕒 Yesterday

Camp Strategy

1 - 10

🏗️ Construction

📦 Logistics

🏨 Hospitality

Principal DevOps Engineer leading AWS, Kubernetes, and Terraform infrastructure for Campspot’s campground reservation software and camping marketplace. Driving automation, reliability, security, and cloud-cost optimization.

🕒 2 days ago

Fingerprint

51 - 200

🔒 Cybersecurity

🔌 API

☁️ SaaS

Staff SRE improving reliability for Fingerprint’s device-intelligence platform. Defining SLOs, strengthening incident operations, and enabling safe AI-assisted system operations.

🇺🇸 United States – Remote

💵 $177k - $240k / year

💰 $32M Series B on 2021-11

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)