Network DevOps Engineer, RDMA Fabric Automation

🕒 August 6

🇺🇸 United States – Remote

💵 $90k - $130k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 3%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Vultr

Vultr

201 - 500 employees

Founded 2014

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

💰 $329M Debt Financing - Vultr on 2025-06

Artificial Intelligence • B2B • Hardware

Vultr is a global cloud infrastructure provider offering on-demand virtual machines, bare-metal servers, GPU-accelerated instances, managed databases, object and block storage, Kubernetes, and networking services. The platform emphasizes AI and HPC workloads with a broad selection of AMD and NVIDIA GPUs, fast networking, and 32+ data center regions, plus a marketplace of deployable apps and developer-friendly APIs. Vultr targets developers and businesses seeking affordable, scalable, and compliant cloud compute and storage alternatives to hyperscalers.

📋 Description

• Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers • Build Ansible and Python-based frameworks to provision, validate, and remediate underlay and overlay networks • Integrate network automation with Vultr’s source-of-truth systems, including NetBox and OpsMill, for intent-driven configuration and validation • Develop telemetry ingestion and correlation pipelines using gNMI, Prometheus, Kafka, and custom collectors • Collaborate with platform, orchestration, and product engineering teams to optimize RDMA performance, PFC/ECN behavior, and path symmetry across fabrics • Implement CI/CD workflows for network configuration changes, including validation, pre-checks, and rollbacks • Investigate complex network behaviors across flow hashing, congestion domains, ECMP, and overlay interactions • Contribute to next-generation GPU and AI interconnect fabric design and integration into Vultr’s global network architecture

🎯 Requirements

• Solid understanding of modern data center networking: EVPN-VXLAN, BGP, MLAG, QoS, and traffic engineering • Deep familiarity with RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design • Strong experience with automation frameworks such as Ansible • Experience with Python, Golang, Rust, or PHP • Comfort working with telemetry and monitoring stacks such as Prometheus, Grafana, Loki, and ELK • Previous experience integrating with NetBox, Nautobot, OpsMill, or similar topology and configuration source-of-truth systems • Familiarity with CI/CD systems such as GitHub Actions, Jenkins, and ArgoCD • Strong Linux networking background, including namespaces, netlink, and system-level debugging • Legally authorized to work in the United States

🏖️ Benefits

• 100% company-paid insurance premiums for employee medical, dental and vision plans • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Increased PTO at 3 year and 10 year anniversary • 1 month paid sabbatical every 5 years • Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription

Apply Now

Similar Jobs

🕒 August 5

MeridianLink

501 - 1000

💳 Fintech

🏦 Banking

☁️ SaaS

Senior SRE owning reliability, observability, and scalability for MeridianLink’s financial SaaS applications. Designing resilient AWS/Azure infrastructure, automation, incident response, and security practices.

🇺🇸 United States – Remote

💵 $104.1k - $177.6k / year

💰 $485M Post-IPO Debt on 2021-11

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

🕒 August 5

NEC Software Solutions

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior DevOps Engineer managing AWS cloud infrastructure, Kubernetes, Terraform, and CI/CD for NEC SWS public-sector systems. Hybrid role requiring 50% office attendance and SC eligibility.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 5

ARUP Laboratories

1001 - 5000

💼 Consulting

🏥 Healthcare

🧬 Biotechnology

DevOps Engineer III building secure, scalable cloud platforms for ARUP Laboratories’ clinical genomics systems. Automating releases, infrastructure, and developer workflows across AWS and enterprise technologies.

Ansible

AWS

Cloud

Docker

Linux

MongoDB

RabbitMQ

SaltStack

Shell Scripting

Splunk

SQL

🕒 August 5

Eliza

11 - 50

🤖 Artificial Intelligence

💼 Consulting

🏢 Enterprise

AI services company engineer building industry-specific ChatGPT experiences and AI prototypes for customers. Leading executive workshops and partnering with OpenAI to turn opportunities into deployable solutions.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 5

Sporttrade

11 - 50

💼 Consulting

📣 Marketing

🎲 Gambling

Site Reliability Engineer operating Sporttrade’s regulated sports betting exchange across cloud and datacenter infrastructure. Automating operations, improving observability, and leading incident response for a live marketplace.

🇺🇸 United States – Remote

💵 $150k - $170k / year

💰 $36M Funding Round on 2021-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)