
1001 - 5000 employés
🤖 Intelligence artificielle
🏢 Entreprise
☁️ SaaS
Artificial Intelligence • Enterprise • SaaS
Le groupe Nebius construit l’une des principales entreprises d’infrastructure AI au monde, en se concentrant sur la fourniture de la puissance de calcul, du stockage et des outils nécessaires aux développeurs dans le domaine de l’AI. Basée en Europe et cotée au Nasdaq, Nebius dispose d’une présence mondiale avec des centres de R&D en Europe, en Amérique du Nord et en Israël. L’offre principale de l’entreprise est une plateforme cloud centrée sur l’AI, conçue pour des workloads AI intensifs, complétée par diverses autres activités impliquées dans le développement de l’AI générative, l’edtech et les technologies autonomes.
Offre fantôme probable
🕒 il y a 6 mois
🗣️🇺🇸🇬🇧 Anglais requis
Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

1001 - 5000 employés
🤖 Intelligence artificielle
🏢 Entreprise
☁️ SaaS
Artificial Intelligence • Enterprise • SaaS
Le groupe Nebius construit l’une des principales entreprises d’infrastructure AI au monde, en se concentrant sur la fourniture de la puissance de calcul, du stockage et des outils nécessaires aux développeurs dans le domaine de l’AI. Basée en Europe et cotée au Nasdaq, Nebius dispose d’une présence mondiale avec des centres de R&D en Europe, en Amérique du Nord et en Israël. L’offre principale de l’entreprise est une plateforme cloud centrée sur l’AI, conçue pour des workloads AI intensifs, complétée par diverses autres activités impliquées dans le développement de l’AI générative, l’edtech et les technologies autonomes.
• Tune the performance of GPU clusters and InfiniBand networks for optimal operation in HPC and GPU-based environments • Analyze and troubleshoot root causes of issues related to GPUs and InfiniBand networks and propose corrective actions • Integrate new hardware into existing infrastructure, including support for new GPU hardware through Kubernetes, QEMU, and KVM software stacks • Enhance automation systems for proactive monitoring, issue detection, and resolution in GPU and InfiniBand environments • Configure and manage GPU devices and InfiniBand fabrics for efficient and reliable operation • Work with hardware virtualization and device emulation technologies in multi-GPU, HPC environments • Analyze, troubleshoot, and improve infrastructure to support new hardware and fine-tune system performance
• 5+ years of professional experience in system-level software development, focused on performance optimization and low-level programming • 3+ years of hands-on experience with Linux systems, including administration, troubleshooting, and performance tuning • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing systems • Strong proficiency in one or more performance-oriented programming languages: C/C++, Go, or Python • Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking is a plus • Proven track record of analyzing and optimizing HPC workloads is a plus • Familiarity with RDMA, RoCE, and InfiniBand protocols is a plus • Background in Software-Defined Networking and experience with HPC cluster networking is a plus • Understanding of QEMU/KVM virtualization and managing virtualized environments is a plus • Experience with deep learning frameworks such as PyTorch and TensorFlow and their integration with HPC systems is a plus • Familiarity with collective communication libraries such as MPI and NCCL is a plus • Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility • Must complete coding interviews as part of the process
• Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams • Fast moving • Bold thinking • Constant growth • Meaningful impact • Trust and real ownership • Opportunity to shape the future of AI • Equal opportunity and inclusive workplace
Postuler Maintenant