
201 - 500 employees
Founded 2014
đ¤ Artificial Intelligence
đ¤ B2B
đ§ Hardware
đ° $329M Debt Financing - Vultr on 2025-06
Artificial Intelligence ⢠B2B ⢠Hardware
Vultr is a global cloud infrastructure provider offering on-demand virtual machines, bare-metal servers, GPU-accelerated instances, managed databases, object and block storage, Kubernetes, and networking services. The platform emphasizes AI and HPC workloads with a broad selection of AMD and NVIDIA GPUs, fast networking, and 32+ data center regions, plus a marketplace of deployable apps and developer-friendly APIs. Vultr targets developers and businesses seeking affordable, scalable, and compliant cloud compute and storage alternatives to hyperscalers.
đ 5 days ago
đşđ¸ United States â Remote
đľ $140k - $160k / year
â° Full Time
đ Senior
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 3%
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
Founded 2014
đ¤ Artificial Intelligence
đ¤ B2B
đ§ Hardware
đ° $329M Debt Financing - Vultr on 2025-06
Artificial Intelligence ⢠B2B ⢠Hardware
Vultr is a global cloud infrastructure provider offering on-demand virtual machines, bare-metal servers, GPU-accelerated instances, managed databases, object and block storage, Kubernetes, and networking services. The platform emphasizes AI and HPC workloads with a broad selection of AMD and NVIDIA GPUs, fast networking, and 32+ data center regions, plus a marketplace of deployable apps and developer-friendly APIs. Vultr targets developers and businesses seeking affordable, scalable, and compliant cloud compute and storage alternatives to hyperscalers.
⢠Build and lead the Production Site Reliability Engineering team ⢠Hire, mentor, and grow SREs and database engineers responsible for the production control plane ⢠Own availability, performance, and operability of the production web stack ⢠Lead incident response for production-impacting events ⢠Drive postmortems and convert lessons into system improvements, runbooks, and alerting ⢠Own configuration management for the production environment ⢠Partner with engineering teams to ensure new services are observable, deployable, and documented before reaching customers ⢠Drive control plane re-architecture, including new infrastructure, parallel runs, cutover, and production-readiness validation ⢠Set the operational roadmap covering capacity planning, scaling, disaster recovery, and architectural evolution ⢠Establish and maintain monitoring, alerting, and observability across the production stack ⢠Manage production database and caching environments, including replication topology, performance tuning, backup verification, and failover testing ⢠Foster operational excellence through operational reviews, SLO definition and tracking, blameless postmortems, and continuous improvement of runbooks and deployment processes
⢠10+ years of professional experience in site reliability engineering or infrastructure operations, with at least 2 years in a team lead or management role ⢠Deep experience operating production web stacks at scale ⢠Comfortable debugging performance issues across the full request path from load balancer to database ⢠Strong Linux systems knowledge: networking, systemd, package management, firewall configuration, and performance tuning ⢠Track record of building and operating monitoring and alerting systems and driving observability and proactive incident response ⢠Experience with configuration management at scale, such as Puppet, Ansible, Chef, or similar ⢠Strong written and verbal communication for runbooks, postmortems, incident communications, and cross-functional coordination ⢠Experience hiring and building engineering teams ⢠Experience with database replication, backup strategies, and failure modes ⢠Experience leading production migrations or major infrastructure transitions ⢠Bonus: experience with Harvester or similar HCI platforms for running production workloads on VMs ⢠Bonus: experience with Redis cluster architecture, failover, and performance tuning at scale ⢠Bonus: experience with HAProxy configuration, keepalived, and load balancing strategies ⢠Bonus: experience with PCI compliance environments and systems handling cardholder data ⢠Bonus: experience writing PHP ⢠Bonus: experience with Cloudflare or similar CDN/edge platforms and DNS cutover ⢠No specific educational credential stated
⢠100% company-paid insurance premiums for employee medical, dental and vision plans ⢠401(k) plan that matches 100% up to 4%, with immediate vesting ⢠Professional Development Reimbursement of $2,500 each year ⢠11 Holidays + Paid Time Off Accrual + Rollover Plan ⢠Increased PTO at 3 year and 10 year anniversary ⢠1 month paid sabbatical every 5 years ⢠Anniversary Bonus each year ⢠$500 stipend for remote office setup in first year + $400 each following year ⢠Internet reimbursement up to $75 per month ⢠Gym membership reimbursement up to $50 per month ⢠Company paid Wellable subscription
Apply Nowđ 5 days ago
Senior Site Reliability Engineer operating Blackpoint Cyberâs cybersecurity infrastructure. Automating AWS, Kubernetes, CI/CD, streaming, observability, and incident response for reliable systems.
đşđ¸ United States â Remote
đľ $150k - $187k / year
đ° $190M Series C on 2023-06
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ 5 days ago
Platform SRE maintaining secure cloud infrastructure, CI/CD, and observability for First Dueâs fire and EMS software. Improving reliability, deployments, scalability, and incident response.
đşđ¸ United States â Remote
đľ $165k / year
đ° $355M Private Equity Round - First Due on 2025-08
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ 5 days ago
DevOps Engineer building Kubernetes, cloud, and deployment infrastructure for Emaâs agentic AI platform. Ensuring secure, observable, reliable scaling for enterprise workflows.
đşđ¸ United States â Remote
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ September 26
Lead DevOps Engineer operating multi-account AWS and Kubernetes platforms for ICF, a global advisory and technology services provider. Automating secure, observable enterprise application delivery.
đşđ¸ United States â Remote
đľ $131.3k - $223.1k / year
đ° $29M Grant on 2023-03
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ September 26
DevSecOps Security Engineer securing Mastercamâs cloud-native manufacturing software, infrastructure, and AI platforms. Integrating Kubernetes, CI/CD, identity, and application security controls.
đşđ¸ United States â Remote
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)