Senior HPC Cluster Engineer

Vermutlich ein Geisterjob

🕒 vor 5 Monaten

🌐 Spanien, Vereinigtes Königreich – Remote

infoinfo

⏰ Vollzeit

🟠 Senior

👷🏻‍♀️ Ingenieur

👻 Geisterscore 70%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung ßberprßfen.

Logo of Nebius Group

Nebius Group

1001 - 5000 Mitarbeiter

🤖 Künstliche Intelligenz

🏢 Unternehmen

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Die Nebius Group baut eines der weltweit fßhrenden Unternehmen fßr KI-Infrastruktur auf und konzentriert sich darauf, die notwendige Rechenleistung, Speicherkapazität und Tools fßr Entwickler im KI-Bereich bereitzustellen. Mit Sitz in Europa und an der Nasdaq notiert verfßgt Nebius ßber eine globale Präsenz mit F&E-Zentren in Europa, Nordamerika und Israel. Das zentrale Angebot des Unternehmens ist eine KI-zentrierte Cloud-Plattform, die fßr rechenintensive KI-Workloads ausgelegt ist, ergänzt durch verschiedene weitere Geschäftsbereiche in den Bereichen Generative KI, Edtech und autonome Technologien.

Beschreibung

• Tune the performance of GPU clusters and InfiniBand networks for optimal operation in HPC and GPU-based environments • Analyze and troubleshoot root causes of issues related to GPUs and InfiniBand networks and propose corrective actions • Integrate new hardware into existing infrastructure, including support for new GPU hardware through Kubernetes, QEMU, and KVM software stacks • Enhance automation systems for proactive monitoring, issue detection, and resolution in GPU and InfiniBand environments • Configure and manage GPU devices and InfiniBand fabrics for efficient and reliable operation • Work with hardware virtualization and device emulation technologies in multi-GPU, HPC environments • Analyze, troubleshoot, and improve infrastructure to support new hardware and fine-tune system performance

🎯 Anforderungen

• 5+ years of professional experience in system-level software development, focused on performance optimization and low-level programming • 3+ years of hands-on experience with Linux systems, including administration, troubleshooting, and performance tuning • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing systems • Strong proficiency in one or more performance-oriented programming languages: C/C++, Go, or Python • Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking is a plus • Proven track record of analyzing and optimizing HPC workloads is a plus • Familiarity with RDMA, RoCE, and InfiniBand protocols is a plus • Background in Software-Defined Networking and experience with HPC cluster networking is a plus • Understanding of QEMU/KVM virtualization and managing virtualized environments is a plus • Experience with deep learning frameworks such as PyTorch and TensorFlow and their integration with HPC systems is a plus • Familiarity with collective communication libraries such as MPI and NCCL is a plus • Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility • Must complete coding interviews as part of the process

🏖️ Vorteile

• Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams • Fast moving • Bold thinking • Constant growth • Meaningful impact • Trust and real ownership • Opportunity to shape the future of AI • Equal opportunity and inclusive workplace

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 6 Monaten

ElevenLabs

1 - 10

🤖 Künstliche Intelligenz

📱 Medien

Forward Deployed Engineer Strategist working with a creative team to deliver innovative voice AI solutions for clients worldwide. Engaging with customers to identify needs and implementing tailored integrations.

🇪🇸 Spanien – Remote

💰 €19.000.000 Series A im 2023-06

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich