Senior HPC Cluster Engineer

Likely ghost job

🕒 April 2

🌐 Spain, United Kingdom – Remote

infoinfo

⏰ Full Time

🟠 Senior

👷🏻‍♀️ Engineer

👻 Ghost score 70%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Nebius Group

Nebius Group

1001 - 5000 employees

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Nebius Group is building one of the world’s leading AI infrastructure companies, focusing on providing the necessary compute, storage, and tools for developers in the AI space. Based in Europe and listed on Nasdaq, Nebius has a global presence with R&D centers across Europe, North America, and Israel. The company's primary offering is an AI-centric cloud platform designed for intensive AI workloads, complemented by various other businesses involved in generative AI development, edtech, and autonomous technology.

📋 Description

• Tune the performance of GPU clusters and InfiniBand networks for optimal operation in HPC and GPU-based environments • Analyze and troubleshoot root causes of issues related to GPUs and InfiniBand networks and propose corrective actions • Integrate new hardware into existing infrastructure, including support for new GPU hardware through Kubernetes, QEMU, and KVM software stacks • Enhance automation systems for proactive monitoring, issue detection, and resolution in GPU and InfiniBand environments • Configure and manage GPU devices and InfiniBand fabrics for efficient and reliable operation • Work with hardware virtualization and device emulation technologies in multi-GPU, HPC environments • Analyze, troubleshoot, and improve infrastructure to support new hardware and fine-tune system performance

🎯 Requirements

• 5+ years of professional experience in system-level software development, focused on performance optimization and low-level programming • 3+ years of hands-on experience with Linux systems, including administration, troubleshooting, and performance tuning • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing systems • Strong proficiency in one or more performance-oriented programming languages: C/C++, Go, or Python • Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking is a plus • Proven track record of analyzing and optimizing HPC workloads is a plus • Familiarity with RDMA, RoCE, and InfiniBand protocols is a plus • Background in Software-Defined Networking and experience with HPC cluster networking is a plus • Understanding of QEMU/KVM virtualization and managing virtualized environments is a plus • Experience with deep learning frameworks such as PyTorch and TensorFlow and their integration with HPC systems is a plus • Familiarity with collective communication libraries such as MPI and NCCL is a plus • Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility • Must complete coding interviews as part of the process

🏖️ Benefits

• Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams • Fast moving • Bold thinking • Constant growth • Meaningful impact • Trust and real ownership • Opportunity to shape the future of AI • Equal opportunity and inclusive workplace

Apply Now

Similar Jobs

🕒 March 5

ElevenLabs

1 - 10

🤖 Artificial Intelligence

📱 Media

Forward Deployed Engineer Strategist working with a creative team to deliver innovative voice AI solutions for clients worldwide. Engaging with customers to identify needs and implementing tailored integrations.

Python