Staff Site Reliability Engineer

🔥 0 minutes ago

🏄 California – Remote

info

💵 $240k - $300k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Skydio

Skydio

501 - 1000 employees

Founded 2014

🎖️ Defense

🏭 Manufacturing

📦 Logistics

💰 $170M Series E - Skydio on 2024-11

Defense • Manufacturing • Logistics

Skydio is a manufacturer of autonomous drone systems that combine advanced onboard artificial intelligence, sensors, and cloud software to automate flight for inspection, security, public safety, and mapping missions. Its product family (including the X10, X10D, R10, docking stations, and software such as Skydio Autonomy and DFR Command) is designed for rapid situational awareness, automated inspections, site security, and Drone-as-First-Responder deployments for government, utilities, and enterprise customers. Skydio delivers integrated hardware, software, and services to enable safer, faster, and more autonomous operations in challenging environments.

📋 Description

• Build, operate, and troubleshoot production Kubernetes/EKS clusters • Perform Kubernetes upgrades, node rollouts, and cluster maintenance • Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage • Define and maintain infrastructure using Terraform • Build and operate CI/CD and deployment infrastructure • Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases • Build monitoring, alerting, and observability for critical infrastructure • Participate in on-call rotations and respond to production incidents • Identify and solve infrastructure scaling and reliability problems • Automate operational work using Python, Go, or similar languages • Help expand infrastructure across new regions and deployment environments

🎯 Requirements

• 8+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps, Production Engineer, or equivalent infrastructure role • Strong hands-on experience operating Kubernetes • Experience managing Kubernetes/EKS upgrades and production clusters • Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases • Production experience with Terraform or similar infrastructure-as-code tooling • Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins • Experience diagnosing production infrastructure and networking problems • Experience solving meaningful scaling or reliability challenges • U.S. person status verification and ability to access controlled or restricted information as required • Bonus: Helm and GitOps experience • Bonus: Datadog or similar observability tooling • Bonus: PostgreSQL/database operations experience • Bonus: Multi-region infrastructure experience • Bonus: On-premises or disconnected deployment experience • Bonus: Streaming or high-throughput distributed systems experience

🏖️ Benefits

• Competitive base salary • Equity in the form of stock options • Comprehensive benefits package • Relocation assistance may be provided for eligible roles • Company group health insurance plans • Paid vacation time • Sick leave • Holiday pay • 401K savings plan

Apply Now

Similar Jobs

🔥 2 hours ago

Renesas Electronics

10,000+ employees

🏭 Manufacturing

🏥 Healthcare

📦 Logistics

Quality and reliability engineer qualifying AI server power modules at Renesas, a global semiconductor solutions company. Developing reliability tests, component qualification processes, validation plans, and manufacturing prototypes for high-performance computing products.

🔥 6 hours ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

GDIT DevSecOps Engineer maintaining secure AWS, Kubernetes, and GitLab pipelines. Supporting defense and intelligence technology, video encoding, automated security scanning, and system performance validation.

🕒 Yesterday

CareSource

1001 - 5000

🏥 Healthcare

🛡️ Insurance

⚕️ Healthcare Insurance

Cloud, DevOps, and SRE architect designing infrastructure, deployment automation, and reliability strategies for CareSource’s digital healthcare platform. Scaling enterprise technology for national delivery.

🕒 Yesterday

AlphaSense

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

Principal engineer architecting scalable CI/CD and Kubernetes delivery systems for AlphaSense, an AI-powered market intelligence platform. Defining progressive delivery standards and developer self-service across global engineering teams.

🕒 Yesterday

Honeycomb.io

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Staff Field Reliability Engineer architecting managed observability infrastructure and resolving high-stakes customer escalations for Honeycomb.io. Leading AWS, Kubernetes, OpenTelemetry, incident response, and strategic technical initiatives.