
201 - 500 funcionários
Fundada em 2014
🤖 Inteligência Artificial
🤝 B2B
🔧 Hardware
💰 $329.000.000 Debt Financing - Vultr em 2025-06
Artificial Intelligence • B2B • Hardware
A Vultr é um provedor global de infraestrutura em nuvem que oferece máquinas virtuais sob demanda, servidores bare-metal, instâncias aceleradas por GPU, bancos de dados gerenciados, armazenamento de objetos e em blocos, Kubernetes e serviços de rede. A plataforma enfatiza cargas de trabalho de IA e HPC com uma ampla seleção de GPUs AMD e NVIDIA, rede rápida e mais de 32 regiões de data centers, além de um marketplace de aplicativos implantáveis e APIs amigáveis para desenvolvedores. A Vultr tem como público-alvo desenvolvedores e empresas que buscam alternativas de computação e armazenamento em nuvem acessíveis, escaláveis e em conformidade com os regulamentos em relação aos hyperscalers.
🕒 2 dias atrás
🇺🇸 Estados Unidos – Remoto (EUA)
💵 $140.000 - $160.000 / ano
⏰ Tempo Integral
🟠 Sênior
🔴 Especialista
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
👻 Score fantasma 2%
🗣️🇺🇸🇬🇧 Inglês obrigatório
Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

201 - 500 funcionários
Fundada em 2014
🤖 Inteligência Artificial
🤝 B2B
🔧 Hardware
💰 $329.000.000 Debt Financing - Vultr em 2025-06
Artificial Intelligence • B2B • Hardware
A Vultr é um provedor global de infraestrutura em nuvem que oferece máquinas virtuais sob demanda, servidores bare-metal, instâncias aceleradas por GPU, bancos de dados gerenciados, armazenamento de objetos e em blocos, Kubernetes e serviços de rede. A plataforma enfatiza cargas de trabalho de IA e HPC com uma ampla seleção de GPUs AMD e NVIDIA, rede rápida e mais de 32 regiões de data centers, além de um marketplace de aplicativos implantáveis e APIs amigáveis para desenvolvedores. A Vultr tem como público-alvo desenvolvedores e empresas que buscam alternativas de computação e armazenamento em nuvem acessíveis, escaláveis e em conformidade com os regulamentos em relação aos hyperscalers.
• Build and lead the Production Site Reliability Engineering team • Hire, mentor, and grow SREs and database engineers responsible for the production control plane • Own availability, performance, and operability of the production web stack • Lead incident response for production-impacting events • Drive postmortems and convert lessons into system improvements, runbooks, and alerting • Own configuration management for the production environment • Partner with engineering teams to ensure new services are observable, deployable, and documented before reaching customers • Drive control plane re-architecture, including new infrastructure, parallel runs, cutover, and production-readiness validation • Set the operational roadmap covering capacity planning, scaling, disaster recovery, and architectural evolution • Establish and maintain monitoring, alerting, and observability across the production stack • Manage production database and caching environments, including replication topology, performance tuning, backup verification, and failover testing • Foster operational excellence through operational reviews, SLO definition and tracking, blameless postmortems, and continuous improvement of runbooks and deployment processes
• 10+ years of professional experience in site reliability engineering or infrastructure operations, with at least 2 years in a team lead or management role • Deep experience operating production web stacks at scale • Comfortable debugging performance issues across the full request path from load balancer to database • Strong Linux systems knowledge: networking, systemd, package management, firewall configuration, and performance tuning • Track record of building and operating monitoring and alerting systems and driving observability and proactive incident response • Experience with configuration management at scale, such as Puppet, Ansible, Chef, or similar • Strong written and verbal communication for runbooks, postmortems, incident communications, and cross-functional coordination • Experience hiring and building engineering teams • Experience with database replication, backup strategies, and failure modes • Experience leading production migrations or major infrastructure transitions • Bonus: experience with Harvester or similar HCI platforms for running production workloads on VMs • Bonus: experience with Redis cluster architecture, failover, and performance tuning at scale • Bonus: experience with HAProxy configuration, keepalived, and load balancing strategies • Bonus: experience with PCI compliance environments and systems handling cardholder data • Bonus: experience writing PHP • Bonus: experience with Cloudflare or similar CDN/edge platforms and DNS cutover • No specific educational credential stated
• 100% company-paid insurance premiums for employee medical, dental and vision plans • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Increased PTO at 3 year and 10 year anniversary • 1 month paid sabbatical every 5 years • Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription
Candidatar-se🕒 2 dias atrás
Senior DevOps Engineer owning cloud, on-premises, and GPU infrastructure for a defense-tech AI company. Building Dockerized deployments, CI/CD, observability, and backend APIs.
🇺🇸 Estados Unidos – Remoto (EUA)
⏰ Tempo Integral
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 2 dias atrás
Senior Site Reliability Engineer operating Blackpoint Cyber’s cybersecurity infrastructure. Automating AWS, Kubernetes, CI/CD, streaming, observability, and incident response for reliable systems.
🇺🇸 Estados Unidos – Remoto (EUA)
💵 $150.000 - $187.000 / ano
💰 $190.000.000 Series C em 2023-06
⏰ Tempo Integral
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 2 dias atrás
Platform SRE maintaining secure cloud infrastructure, CI/CD, and observability for First Due’s fire and EMS software. Improving reliability, deployments, scalability, and incident response.
🇺🇸 Estados Unidos – Remoto (EUA)
💵 $165.000 / ano
💰 $355.000.000 Private Equity Round - First Due em 2025-08
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 2 dias atrás
DevOps Engineer building Kubernetes, cloud, and deployment infrastructure for Ema’s agentic AI platform. Ensuring secure, observable, reliable scaling for enterprise workflows.
🇺🇸 Estados Unidos – Remoto (EUA)
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 3 dias atrás
Principal DevSecOps Engineer securing GDIT government and defense cloud infrastructure, pipelines, and AI workloads. Automating compliance, Kubernetes, and Infrastructure as Code operations.
🇺🇸 Estados Unidos – Remoto (EUA)
💵 $136.000 - $184.000 / ano
⏰ Tempo Integral
🔴 Especialista
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🦅 Patrocina Visto H1B
🗣️🇺🇸🇬🇧 Inglês obrigatório