Network DevOps Engineer, RDMA Fabric Automation

🕒 vor 29 Tagen

🇺🇸 Vereinigte Staaten – Remote

💵 $90.000 - $130.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 2%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Vultr

Vultr

201 - 500 Mitarbeiter

Gegründet 2014

🤖 Künstliche Intelligenz

🤝 B2B

🔧 Hardware

💰 €329.000.000 Debt Financing - Vultr im 2025-06

Artificial Intelligence • B2B • Hardware

Vultr ist ein globaler Anbieter von Cloud-Infrastrukturen, der bedarfsgesteuerte virtuelle Maschinen, Bare-Metal-Server, GPU-beschleunigte Instanzen, verwaltete Datenbanken, Objekt- und Blockspeicher, Kubernetes- und Netzwerklösungen anbietet. Die Plattform legt den Schwerpunkt auf KI- und HPC-Workloads mit einer breiten Auswahl an AMD- und NVIDIA-GPUs, schnellen Netzwerken und über 32 Datenzentrumsregionen sowie einem Marktplatz für bereitstellbare Apps und entwicklerfreundliche APIs. Vultr richtet sich an Entwickler und Unternehmen, die nach kostengünstigen, skalierbaren und konformen Alternativen zu Hyperscalern für Cloud-Computing und Speicherlösungen suchen.

Beschreibung

• Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers • Build Ansible and Python-based frameworks to provision, validate, and remediate underlay and overlay networks • Integrate network automation with Vultr’s source-of-truth systems, including NetBox and OpsMill, for intent-driven configuration and validation • Develop telemetry ingestion and correlation pipelines using gNMI, Prometheus, Kafka, and custom collectors • Collaborate with platform, orchestration, and product engineering teams to optimize RDMA performance, PFC/ECN behavior, and path symmetry across fabrics • Implement CI/CD workflows for network configuration changes, including validation, pre-checks, and rollbacks • Investigate complex network behaviors across flow hashing, congestion domains, ECMP, and overlay interactions • Contribute to next-generation GPU and AI interconnect fabric design and integration into Vultr’s global network architecture

🎯 Anforderungen

• Solid understanding of modern data center networking: EVPN-VXLAN, BGP, MLAG, QoS, and traffic engineering • Deep familiarity with RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design • Strong experience with automation frameworks such as Ansible • Experience with Python, Golang, Rust, or PHP • Comfort working with telemetry and monitoring stacks such as Prometheus, Grafana, Loki, and ELK • Previous experience integrating with NetBox, Nautobot, OpsMill, or similar topology and configuration source-of-truth systems • Familiarity with CI/CD systems such as GitHub Actions, Jenkins, and ArgoCD • Strong Linux networking background, including namespaces, netlink, and system-level debugging • Legally authorized to work in the United States

🏖️ Vorteile

• 100% company-paid insurance premiums for employee medical, dental and vision plans • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Increased PTO at 3 year and 10 year anniversary • 1 month paid sabbatical every 5 years • Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 29 Tagen

Capgemini

10.000+ Mitarbeiter

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

Enterprise AI practice leader at Capgemini, a global technology transformation partner. Driving AI strategy, governance, delivery, growth, and organizational transformation for enterprise clients.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 29 Tagen

CampMinder

51 - 200

💼 Beratung

🏥 Gesundheitswesen

🏨 Gastgewerbe

Senior DevSecOps Engineer securing Campminder’s software for summer camps. Hardening cloud infrastructure, embedding security in CI/CD, and leading compliance and threat-remediation work.

🇺🇸 Vereinigte Staaten – Remote

💵 $180.000 - $200.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

Azure

Cloud

Cyber Security

Kubernetes

Vault

🕒 vor 1 Monat

MeridianLink

501 - 1000

💳 Fintech

🏦 Bankwesen

☁️ SaaS

Senior SRE owning reliability, observability, and scalability for MeridianLink’s financial SaaS applications. Designing resilient AWS/Azure infrastructure, automation, incident response, and security practices.

🇺🇸 Vereinigte Staaten – Remote

💵 $104.148 - $177.600 / Jahr

💰 €485.000.000 Post-IPO Debt im 2021-11

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

eFinancial

201 - 500

🏥 Gesundheitswesen

💼 Beratung

🛡️ Versicherung

DevOps Team Lead building AWS infrastructure, CI/CD pipelines, and shared engineering platforms for a life insurance provider. Coaching engineers and improving deployment reliability and automation.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

NEC Software Solutions

5001 - 10000

🏥 Gesundheitswesen

💼 Beratung

📦 Logistik

Senior DevOps Engineer managing AWS cloud infrastructure, Kubernetes, Terraform, and CI/CD for NEC SWS public-sector systems. Hybrid role requiring 50% office attendance and SC eligibility.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich