Network DevOps Engineer, RDMA Fabric Automation

🕒 Agosto 6

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $90.000 - $130.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 3%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Vultr

Vultr

201 - 500 funcionários

Fundada em 2014

🤖 Inteligência Artificial

🤝 B2B

🔧 Hardware

💰 $329.000.000 Debt Financing - Vultr em 2025-06

Artificial Intelligence • B2B • Hardware

A Vultr é um provedor global de infraestrutura em nuvem que oferece máquinas virtuais sob demanda, servidores bare-metal, instâncias aceleradas por GPU, bancos de dados gerenciados, armazenamento de objetos e em blocos, Kubernetes e serviços de rede. A plataforma enfatiza cargas de trabalho de IA e HPC com uma ampla seleção de GPUs AMD e NVIDIA, rede rápida e mais de 32 regiões de data centers, além de um marketplace de aplicativos implantáveis e APIs amigáveis para desenvolvedores. A Vultr tem como público-alvo desenvolvedores e empresas que buscam alternativas de computação e armazenamento em nuvem acessíveis, escaláveis e em conformidade com os regulamentos em relação aos hyperscalers.

Descrição

• Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers • Build Ansible and Python-based frameworks to provision, validate, and remediate underlay and overlay networks • Integrate network automation with Vultr’s source-of-truth systems, including NetBox and OpsMill, for intent-driven configuration and validation • Develop telemetry ingestion and correlation pipelines using gNMI, Prometheus, Kafka, and custom collectors • Collaborate with platform, orchestration, and product engineering teams to optimize RDMA performance, PFC/ECN behavior, and path symmetry across fabrics • Implement CI/CD workflows for network configuration changes, including validation, pre-checks, and rollbacks • Investigate complex network behaviors across flow hashing, congestion domains, ECMP, and overlay interactions • Contribute to next-generation GPU and AI interconnect fabric design and integration into Vultr’s global network architecture

🎯 Requisitos

• Solid understanding of modern data center networking: EVPN-VXLAN, BGP, MLAG, QoS, and traffic engineering • Deep familiarity with RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design • Strong experience with automation frameworks such as Ansible • Experience with Python, Golang, Rust, or PHP • Comfort working with telemetry and monitoring stacks such as Prometheus, Grafana, Loki, and ELK • Previous experience integrating with NetBox, Nautobot, OpsMill, or similar topology and configuration source-of-truth systems • Familiarity with CI/CD systems such as GitHub Actions, Jenkins, and ArgoCD • Strong Linux networking background, including namespaces, netlink, and system-level debugging • Legally authorized to work in the United States

🏖️ Benefícios

• 100% company-paid insurance premiums for employee medical, dental and vision plans • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Increased PTO at 3 year and 10 year anniversary • 1 month paid sabbatical every 5 years • Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription

Candidatar-se

Vagas Similares

🕒 Agosto 5

MeridianLink

501 - 1000

💳 Fintech

🏦 Bancário

☁️ SaaS

Senior SRE owning reliability, observability, and scalability for MeridianLink’s financial SaaS applications. Designing resilient AWS/Azure infrastructure, automation, incident response, and security practices.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $104.148 - $177.600 / ano

💰 $485.000.000 Post-IPO Debt em 2021-11

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 5

NEC Software Solutions

5001 - 10000

🏥 Saúde

💼 Consultoria

📦 Logística

Senior DevOps Engineer managing AWS cloud infrastructure, Kubernetes, Terraform, and CI/CD for NEC SWS public-sector systems. Hybrid role requiring 50% office attendance and SC eligibility.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 5

ARUP Laboratories

1001 - 5000

💼 Consultoria

🏥 Saúde

🧬 Biotecnologia

DevOps Engineer III building secure, scalable cloud platforms for ARUP Laboratories’ clinical genomics systems. Automating releases, infrastructure, and developer workflows across AWS and enterprise technologies.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 5

Eliza

11 - 50

🤖 Inteligência Artificial

💼 Consultoria

🏢 Corporativo

AI services company engineer building industry-specific ChatGPT experiences and AI prototypes for customers. Leading executive workshops and partnering with OpenAI to turn opportunities into deployable solutions.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 5

Sporttrade

11 - 50

💼 Consultoria

📣 Marketing

🎲 Jogos de Azar

Site Reliability Engineer operating Sporttrade’s regulated sports betting exchange across cloud and datacenter infrastructure. Automating operations, improving observability, and leading incident response for a live marketplace.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $170.000 / ano

💰 $36.000.000 Funding Round em 2021-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório