Manager, Production Site Reliability Engineering

🔥 12 hours ago

🇺🇸 United States – Remote

💵 $140k - $160k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 3%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Vultr

Vultr

201 - 500 employees

Founded 2014

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

💰 $329M Debt Financing - Vultr on 2025-06

Artificial Intelligence • B2B • Hardware

Vultr is a global cloud infrastructure provider offering on-demand virtual machines, bare-metal servers, GPU-accelerated instances, managed databases, object and block storage, Kubernetes, and networking services. The platform emphasizes AI and HPC workloads with a broad selection of AMD and NVIDIA GPUs, fast networking, and 32+ data center regions, plus a marketplace of deployable apps and developer-friendly APIs. Vultr targets developers and businesses seeking affordable, scalable, and compliant cloud compute and storage alternatives to hyperscalers.

📋 Description

• Build and lead the Production Site Reliability Engineering team • Hire, mentor, and grow SREs and database engineers responsible for the production control plane • Own availability, performance, and operability of the production web stack • Lead incident response for production-impacting events • Drive postmortems and convert lessons into system improvements, runbooks, and alerting • Own configuration management for the production environment • Partner with engineering teams to ensure new services are observable, deployable, and documented before reaching customers • Drive control plane re-architecture, including new infrastructure, parallel runs, cutover, and production-readiness validation • Set the operational roadmap covering capacity planning, scaling, disaster recovery, and architectural evolution • Establish and maintain monitoring, alerting, and observability across the production stack • Manage production database and caching environments, including replication topology, performance tuning, backup verification, and failover testing • Foster operational excellence through operational reviews, SLO definition and tracking, blameless postmortems, and continuous improvement of runbooks and deployment processes

🎯 Requirements

• 10+ years of professional experience in site reliability engineering or infrastructure operations, with at least 2 years in a team lead or management role • Deep experience operating production web stacks at scale • Comfortable debugging performance issues across the full request path from load balancer to database • Strong Linux systems knowledge: networking, systemd, package management, firewall configuration, and performance tuning • Track record of building and operating monitoring and alerting systems and driving observability and proactive incident response • Experience with configuration management at scale, such as Puppet, Ansible, Chef, or similar • Strong written and verbal communication for runbooks, postmortems, incident communications, and cross-functional coordination • Experience hiring and building engineering teams • Experience with database replication, backup strategies, and failure modes • Experience leading production migrations or major infrastructure transitions • Bonus: experience with Harvester or similar HCI platforms for running production workloads on VMs • Bonus: experience with Redis cluster architecture, failover, and performance tuning at scale • Bonus: experience with HAProxy configuration, keepalived, and load balancing strategies • Bonus: experience with PCI compliance environments and systems handling cardholder data • Bonus: experience writing PHP • Bonus: experience with Cloudflare or similar CDN/edge platforms and DNS cutover • No specific educational credential stated

🏖️ Benefits

• 100% company-paid insurance premiums for employee medical, dental and vision plans • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Increased PTO at 3 year and 10 year anniversary • 1 month paid sabbatical every 5 years • Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription

Apply Now

Similar Jobs

🔥 15 hours ago

Gini-Apps

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Senior DevOps Engineer owning cloud, on-premises, and GPU infrastructure for a defense-tech AI company. Building Dockerized deployments, CI/CD, observability, and backend APIs.

🔥 17 hours ago

Blackpoint Cyber

51 - 200

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

Senior Site Reliability Engineer operating Blackpoint Cyber’s cybersecurity infrastructure. Automating AWS, Kubernetes, CI/CD, streaming, observability, and incident response for reliable systems.

🔥 18 hours ago

First Due

201 - 500

☁️ SaaS

🏛️ Government

🤝 B2B

Platform SRE maintaining secure cloud infrastructure, CI/CD, and observability for First Due’s fire and EMS software. Improving reliability, deployments, scalability, and incident response.

🇺🇸 United States – Remote

💵 $165k / year

💰 $355M Private Equity Round - First Due on 2025-08

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 19 hours ago

Fusable

201 - 500

🏗️ Construction

📦 Logistics

🛡️ Insurance

Sr. DevOps Engineer automating AWS infrastructure, containers, and CI/CD pipelines for Fusable’s trucking, agriculture, and construction data solutions. Managing secure, scalable cloud environments and production reliability.

🔥 21 hours ago

Environmental Management Authority

51 - 200

🏥 Healthcare

🛡️ Insurance

💼 Consulting

DevOps Engineer building Kubernetes, cloud, and deployment infrastructure for Ema’s agentic AI platform. Ensuring secure, observable, reliable scaling for enterprise workflows.