Staff AI/ML Infrastructure Engineer

Stelle nicht auf LinkedIn

🕒 vor 4 Monaten

🇺🇸 Vereinigte Staaten – Remote

💵 $145.000 - $160.000 / Jahr

⏰ Vollzeit

🔴 Experte

👷 IT-Infrastrukturingenieur

👻 Geisterscore 25%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Vultr

Vultr

201 - 500 Mitarbeiter

Gegründet 2014

🤖 Künstliche Intelligenz

🤝 B2B

🔧 Hardware

💰 €329.000.000 Debt Financing - Vultr im 2025-06

Artificial Intelligence • B2B • Hardware

Vultr ist ein globaler Anbieter von Cloud-Infrastrukturen, der bedarfsgesteuerte virtuelle Maschinen, Bare-Metal-Server, GPU-beschleunigte Instanzen, verwaltete Datenbanken, Objekt- und Blockspeicher, Kubernetes- und Netzwerklösungen anbietet. Die Plattform legt den Schwerpunkt auf KI- und HPC-Workloads mit einer breiten Auswahl an AMD- und NVIDIA-GPUs, schnellen Netzwerken und über 32 Datenzentrumsregionen sowie einem Marktplatz für bereitstellbare Apps und entwicklerfreundliche APIs. Vultr richtet sich an Entwickler und Unternehmen, die nach kostengünstigen, skalierbaren und konformen Alternativen zu Hyperscalern für Cloud-Computing und Speicherlösungen suchen.

Beschreibung

• Design and maintain GPU and bare metal infrastructure in containerized and physical environments • Build scalable GPU clusters in partnership with networking and provisioning teams • Ensure reliable, high-performance provisioning of GPU infrastructure • Develop automated testing systems for GPU-based platforms • Implement infrastructure solutions for diverse AI/ML workloads • Benchmark, test, and troubleshoot GPU performance at scale • Collaborate with hardware vendors on drivers, firmware, and support • Resolve hardware, software, and performance issues across environments • Optimize rail and cluster performance across architectures • Lead technical direction and mentor engineers on infrastructure best practices

🎯 Anforderungen

• 5+ years experience working with bare metal infrastructure and hardware automation • Hands-on experience with modern NVIDIA/AMD GPU platforms and high-performance networking (RoCE, InfiniBand) • Deep knowledge of BIOS, BMC, firmware, NICs, Redfish/IPMI, and PCIe systems • Strong Linux systems experience including device drivers and package management • Experience building infrastructure automation using Python and Bash • Familiarity with GPU drivers, firmware ecosystems, and vendor collaboration • Experience designing and delivering complex infrastructure products • Proven ability to lead projects and mentor engineers • Experience optimizing multi-cluster GPU environments • Exposure to Machine Learning software stacks and GPU workloads

🏖️ Vorteile

• 100% company-paid insurance premiums for employee medical, dental and vision plans. • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Commitment matters to Vultr! Increased PTO at 3 year and 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 5 Monaten

OpenAI

201 - 500

🤖 Künstliche Intelligenz

☁️ SaaS

🏢 Unternehmen

Principal Software Engineer on Infrastructure Security team at OpenAI safeguarding research and production environments. Leading the development and implementation of planet-scale security systems and primitives.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

OpenAI

201 - 500

🤖 Künstliche Intelligenz

☁️ SaaS

🏢 Unternehmen

Principal Security Engineer at OpenAI protecting critical infrastructure security. Leading security initiatives and ensuring robust security measures across diverse environments.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 6 Monaten

Epistemix

11 - 50

💼 Beratung

🏥 Gesundheitswesen

📣 Marketing

Infrastructure Architect responsible for client integration and deployment of Epistemix's data-driven platform. Requires deep technical expertise in cloud environments and automation tools.

🇺🇸 Vereinigte Staaten – Remote

💰 €7.000.000 Series A - Epistemix im 2024-06

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷 IT-Infrastrukturingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 7 Monaten

Rad AI

51 - 200

🏥 Gesundheitswesen

🏭 Fertigung

🔧 Hardware

Staff Software Engineer designing and operating robust cloud infrastructure for AI-driven healthcare solutions. Collaborating across functions to innovate in a rapidly growing tech environment.

🇺🇸 Vereinigte Staaten – Remote

💵 $175.000 - $230.000 / Jahr

💰 €25.000.000 Series A im 2021-11

⏰ Vollzeit

🔴 Experte

👷 IT-Infrastrukturingenieur

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 9 Monaten

DecisionPoint Corporation

51 - 200

🎖️ Verteidigung

💼 Beratung

📦 Logistik

Zero Trust Infrastructure Engineer implementing PEPs, microsegmentation, and identity-based controls. Supporting federal and DoD-aligned IL5 AWS GovCloud security environments for DecisionPoint.

🗣️🇺🇸🇬🇧 Englisch erforderlich

AWS

Cloud

Cyber Security