GPU Fabric Engineer

🕒 July 13

🇺🇸 United States – Remote

💵 $125k - $135k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Vultr

Vultr

201 - 500 employees

Founded 2014

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

💰 $329M Debt Financing - Vultr on 2025-06

Artificial Intelligence • B2B • Hardware

Vultr is a global cloud infrastructure provider offering on-demand virtual machines, bare-metal servers, GPU-accelerated instances, managed databases, object and block storage, Kubernetes, and networking services. The platform emphasizes AI and HPC workloads with a broad selection of AMD and NVIDIA GPUs, fast networking, and 32+ data center regions, plus a marketplace of deployable apps and developer-friendly APIs. Vultr targets developers and businesses seeking affordable, scalable, and compliant cloud compute and storage alternatives to hyperscalers.

📋 Description

• Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion • Tune fabric performance parameters for distributed AI workloads (NCCL, MPI, collective operations) • Monitor and manage fabrics using NVIDIA UFM (Unified Fabric Manager) for health, topology, and performance visibility • Diagnose and resolve fabric-level issues including link errors, congestion, packet loss, and path asymmetry • Optimize RDMA transport settings, PFC/ECN behavior, and lossless queue configuration for GPU traffic • Validate fabric performance benchmarks and ensure line-rate throughput for AI workloads • Collaborate with GPU Engineers to correlate fabric health with workload performance • Collaborate with networking teams on fabric provisioning, configuration, and remediation • Respond to fabric alerts and degradation events across production GPU clusters • Document fabric troubleshooting procedures, tuning parameters, and validation runbooks

🎯 Requirements

• 3–7 years of experience in network engineering, HPC fabric, or GPU infrastructure • Hands-on experience with InfiniBand and/or RoCE fabrics in GPU cluster environments • Experience with NVIDIA UFM for fabric management, monitoring, and diagnostics • Strong understanding of RDMA transport, lossless Ethernet design, and congestion management (PFC, ECN, DCQCN) • Experience with GPU cluster networking and distributed communication libraries (NCCL, MPI) • Familiarity with GPU platforms and their interconnect requirements (NVIDIA NVLink, NVSwitch, ConnectX) • Experience with fabric diagnostic tools (ibstat, ibqueryerrors, perfquery, etc.) • Proficiency in Python or Bash for scripting and validation • Basic understanding of Linux systems and server hardware • Strong troubleshooting and analytical skills across network and system layers

🏖️ Benefits

• 100% company-paid insurance premiums for employee medical, dental and vision plans. • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Increased PTO at 3 year and 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription

Apply Now

Similar Jobs

🕒 July 13

Terabase Energy

51 - 200

🏗️ Construction

⚡ Energy

RTDS Engineer developing and validating models for utility-scale solar and hybrid projects. Collaborating with engineering teams for Hardware-in-the-Loop simulations and project testing.

🇺🇸 United States – Remote

💵 $140k - $170k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

🦅 H1B Visa Sponsor

info

Firewalls

🕒 July 13

AST SpaceMobile

51 - 200

💼 Consulting

📦 Logistics

✈️ Travel

Design and develop mechanical systems for aerospace applications at AST SpaceMobile. Focus on optimization of mechanical systems from concept through integration and testing.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

Assembly

🕒 July 11

Design Hire LLC

11 - 50

🏗️ Construction

🏭 Manufacturing

💼 Consulting

Piping Engineer with nuclear industry experience supporting piping design, analysis, and engineering in remote role. Focus on ensuring effective and compliant plant operations.

🇺🇸 United States – Remote

💵 $50 - $80 / hour

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

🕒 July 11

ABB

10,000+ employees

💼 Consulting

📦 Logistics

🚗 Transport

Technical lead for engineering activities in protection & control projects at ABB. Overseeing design, documentation, and project management for utility, industrial, and commercial systems.

🇺🇸 United States – Remote

💵 $83.3k - $133.3k / year

💰 $545.9M Post-IPO Debt - ABB on 2023-11

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

🕒 July 11

ProPetro Services, Inc

1001 - 5000

⚡ Energy

☁️ SaaS

Remote Completions Engineer I supports hydraulic fracturing operations from the Remote Operations Center. Responsibilities include monitoring treatment data and maintaining accurate job reporting.

🇺🇸 United States – Remote

💰 Post-IPO Debt on 2023-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer