
201 - 500 employees
Founded 2014
🤖 Artificial Intelligence
🤝 B2B
🔧 Hardware
💰 $329M Debt Financing - Vultr on 2025-06
Artificial Intelligence • B2B • Hardware
Vultr is a global cloud infrastructure provider offering on-demand virtual machines, bare-metal servers, GPU-accelerated instances, managed databases, object and block storage, Kubernetes, and networking services. The platform emphasizes AI and HPC workloads with a broad selection of AMD and NVIDIA GPUs, fast networking, and 32+ data center regions, plus a marketplace of deployable apps and developer-friendly APIs. Vultr targets developers and businesses seeking affordable, scalable, and compliant cloud compute and storage alternatives to hyperscalers.
🕒 4 days ago
🇺🇸 United States – Remote
💵 $90k - $130k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
Founded 2014
🤖 Artificial Intelligence
🤝 B2B
🔧 Hardware
💰 $329M Debt Financing - Vultr on 2025-06
Artificial Intelligence • B2B • Hardware
Vultr is a global cloud infrastructure provider offering on-demand virtual machines, bare-metal servers, GPU-accelerated instances, managed databases, object and block storage, Kubernetes, and networking services. The platform emphasizes AI and HPC workloads with a broad selection of AMD and NVIDIA GPUs, fast networking, and 32+ data center regions, plus a marketplace of deployable apps and developer-friendly APIs. Vultr targets developers and businesses seeking affordable, scalable, and compliant cloud compute and storage alternatives to hyperscalers.
• Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers • Build Ansible and Python-based frameworks to provision, validate, and remediate underlay and overlay networks • Integrate network automation with Vultr’s source-of-truth systems, including NetBox and OpsMill, for intent-driven configuration and validation • Develop telemetry ingestion and correlation pipelines using gNMI, Prometheus, Kafka, and custom collectors • Collaborate with platform, orchestration, and product engineering teams to optimize RDMA performance, PFC/ECN behavior, and path symmetry across fabrics • Implement CI/CD workflows for network configuration changes, including validation, pre-checks, and rollbacks • Investigate complex network behaviors across flow hashing, congestion domains, ECMP, and overlay interactions • Contribute to next-generation GPU and AI interconnect fabric design and integration into Vultr’s global network architecture
• Solid understanding of modern data center networking: EVPN-VXLAN, BGP, MLAG, QoS, and traffic engineering • Deep familiarity with RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design • Strong experience with automation frameworks such as Ansible • Experience with Python, Golang, Rust, or PHP • Comfort working with telemetry and monitoring stacks such as Prometheus, Grafana, Loki, and ELK • Previous experience integrating with NetBox, Nautobot, OpsMill, or similar topology and configuration source-of-truth systems • Familiarity with CI/CD systems such as GitHub Actions, Jenkins, and ArgoCD • Strong Linux networking background, including namespaces, netlink, and system-level debugging • Legally authorized to work in the United States
• 100% company-paid insurance premiums for employee medical, dental and vision plans • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Increased PTO at 3 year and 10 year anniversary • 1 month paid sabbatical every 5 years • Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription
Apply Now🕒 4 days ago
Enterprise AI practice leader at Capgemini, a global technology transformation partner. Driving AI strategy, governance, delivery, growth, and organizational transformation for enterprise clients.
🇺🇸 United States – Remote
💵 $94.2k - $191.9k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 4 days ago
Senior SRE ensuring reliable cloud infrastructure for LeoLabs’ space-security platform. Automating deployments, monitoring systems, responding to incidents, and improving scalability and security.
🇺🇸 United States – Remote
💵 $192k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 4 days ago
Site Reliability Engineer scaling PostHog’s open-source product analytics platform on AWS. Automating Kubernetes infrastructure, multi-account networking, reliability, and incident response.
🕒 4 days ago
Site Reliability Engineer scaling Arango’s cloud-native contextual AI data platform. Automating Kubernetes, AWS, and Google Cloud infrastructure with Golang, observability, and resilient production operations.
🇺🇸 United States – Remote
💰 $27.8M Series B on 2021-10
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 4 days ago
Senior DevSecOps Engineer securing Campminder’s software for summer camps. Hardening cloud infrastructure, embedding security in CI/CD, and leading compliance and threat-remediation work.
🇺🇸 United States – Remote
💵 $180k - $200k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)