ML Infrastructure Engineer

Stelle nicht auf LinkedIn

🕒 vor 3 Monaten

🌐 Vereinigte Staaten, Niederlande – Remote

infoinfo

🏄 California – Remote

infoinfo

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

👷 IT-Infrastrukturingenieur

👻 Geisterscore 42%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Nebius Group

Nebius Group

1001 - 5000 Mitarbeiter

🤖 Künstliche Intelligenz

🏢 Unternehmen

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Die Nebius Group baut eines der weltweit führenden Unternehmen für KI-Infrastruktur auf und konzentriert sich darauf, die notwendige Rechenleistung, Speicherkapazität und Tools für Entwickler im KI-Bereich bereitzustellen. Mit Sitz in Europa und an der Nasdaq notiert verfügt Nebius über eine globale Präsenz mit F&E-Zentren in Europa, Nordamerika und Israel. Das zentrale Angebot des Unternehmens ist eine KI-zentrierte Cloud-Plattform, die für rechenintensive KI-Workloads ausgelegt ist, ergänzt durch verschiedene weitere Geschäftsbereiche in den Bereichen Generative KI, Edtech und autonome Technologien.

Beschreibung

• Work closely with hardware, development teams to profile and analyse GPU performance at the system and kernel level. • Evaluate and compare GPU performance across different platforms, architectures, and software stacks (e.g.,CUDA, ROCm). • Debug and optimise ML workloads to run efficiently on GPU hardware, identifying and resolving performance bottlenecks. • Perform acceptance testing for new GPU clusters, ensuring hardware and software meet performance, stability, and compatibility requirements for AI workloads. • Perform experiments across diverse GPU system configurations to assess the impact of varying interconnect strategies and system-level optimisations on performance and scalability. • Develop tools and dashboards to visualise performance metrics, bottlenecks, and trends. • Contribute to internal tooling, frameworks, and best practices

🎯 Anforderungen

• A profound understanding of theoretical foundations of machine learning • Deep understanding of performance aspects of large neural networks training and inference (data/tensor/context/expert parallelism, offloading, custom kernels, hardware features, attention optimisations, dynamic batching etc.) • Deep experience with modern deep learning frameworks (PyTorch, JAX, Megatron-LM, Tensort-LLM) • Good understanding of the GPU stack: CUDA,NCCL, drivers, and relevant libraries • Familiarity with containerized environments (e.g., Docker, Kubernetes). • Strong communication and ability to work independently

🏖️ Vorteile

• Competitive compensation • Career growth and learning opportunities • Flexibility and work-life balance • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 3 Monaten

Moen

1001 - 5000

🏗️ Bauwesen

🏭 Fertigung

🔧 Hardware

Cloud Infrastructure Architect responsible for designing and deploying cloud solutions at Fortune Brands Innovations. Focused on PaaS offerings and AI integration to transform enterprise IT.

🇺🇸 Vereinigte Staaten – Remote

💵 $105.000 - $165.000 / Jahr

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷 IT-Infrastrukturingenieur

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 3 Monaten

MDaudit

51 - 200

🏥 Gesundheitswesen

⚕️ Krankenversicherung

📋 Compliance

Cloud Infrastructure Engineer responsible for designing, building, securing, and operating cloud infrastructure for healthcare applications across Azure and AWS. Leading migrations and ensuring compliance with security requirements.

🇺🇸 Vereinigte Staaten – Remote

💵 $110.000 - $130.000 / Jahr

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷 IT-Infrastrukturingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

AWS

Azure

Cloud

Firewalls

Linux

RDBMS

SQL

🕒 vor 3 Monaten

Tribe AI

51 - 200

💼 Beratung

🏥 Gesundheitswesen

🤖 Künstliche Intelligenz

Infrastructure Engineer deploying AI systems in client infrastructures. Navigating enterprise constraints and debugging production systems in a variety of environments.

🇺🇸 Vereinigte Staaten – Remote

💵 $200.000 - $300.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

👷 IT-Infrastrukturingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Cribl

501 - 1000

☁️ SaaS

Senior Software Engineer engaged in developing observability data management solutions at Cribl. Contributing to high-quality software deployment and infrastructure improvement in a remote-first environment.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Bayesian Health

11 - 50

🏥 Gesundheitswesen

🤖 Künstliche Intelligenz

⚕️ Krankenversicherung

Infrastructure Engineer building infrastructure and developing CI/CD for clinical AI/ML platform at Bayesian Health. Collaborating with cross-functional teams to ensure reliable and scalable systems.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

👷 IT-Infrastrukturingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich