
1001 - 5000 employees
đ¤ Artificial Intelligence
đ˘ Enterprise
âď¸ SaaS
Artificial Intelligence ⢠Enterprise ⢠SaaS
Nebius Group is building one of the worldâs leading AI infrastructure companies, focusing on providing the necessary compute, storage, and tools for developers in the AI space. Based in Europe and listed on Nasdaq, Nebius has a global presence with R&D centers across Europe, North America, and Israel. The company's primary offering is an AI-centric cloud platform designed for intensive AI workloads, complemented by various other businesses involved in generative AI development, edtech, and autonomous technology.
đ May 12
đ United States, Netherlands â Remote
đ California â Remote
â° Full Time
đĄ Mid-level
đ Senior
đˇ Infrastructure Engineer
đť Ghost score 45%
Improve your chances of getting an interview by checking your resume score before you apply.

1001 - 5000 employees
đ¤ Artificial Intelligence
đ˘ Enterprise
âď¸ SaaS
Artificial Intelligence ⢠Enterprise ⢠SaaS
Nebius Group is building one of the worldâs leading AI infrastructure companies, focusing on providing the necessary compute, storage, and tools for developers in the AI space. Based in Europe and listed on Nasdaq, Nebius has a global presence with R&D centers across Europe, North America, and Israel. The company's primary offering is an AI-centric cloud platform designed for intensive AI workloads, complemented by various other businesses involved in generative AI development, edtech, and autonomous technology.
⢠Work closely with hardware, development teams to profile and analyse GPU performance at the system and kernel level. ⢠Evaluate and compare GPU performance across different platforms, architectures, and software stacks (e.g.,CUDA, ROCm). ⢠Debug and optimise ML workloads to run efficiently on GPU hardware, identifying and resolving performance bottlenecks. ⢠Perform acceptance testing for new GPU clusters, ensuring hardware and software meet performance, stability, and compatibility requirements for AI workloads. ⢠Perform experiments across diverse GPU system configurations to assess the impact of varying interconnect strategies and system-level optimisations on performance and scalability. ⢠Develop tools and dashboards to visualise performance metrics, bottlenecks, and trends. ⢠Contribute to internal tooling, frameworks, and best practices
⢠A profound understanding of theoretical foundations of machine learning ⢠Deep understanding of performance aspects of large neural networks training and inference (data/tensor/context/expert parallelism, offloading, custom kernels, hardware features, attention optimisations, dynamic batching etc.) ⢠Deep experience with modern deep learning frameworks (PyTorch, JAX, Megatron-LM, Tensort-LLM) ⢠Good understanding of the GPU stack: CUDA,NCCL, drivers, and relevant libraries ⢠Familiarity with containerized environments (e.g., Docker, Kubernetes). ⢠Strong communication and ability to work independently
⢠Competitive compensation ⢠Career growth and learning opportunities ⢠Flexibility and work-life balance ⢠Collaborative and innovative culture ⢠Opportunity to work on impactful AI projects ⢠International environment and talented teams
Apply Nowđ May 8
Cloud Infrastructure Engineer responsible for designing, building, securing, and operating cloud infrastructure for healthcare applications across Azure and AWS. Leading migrations and ensuring compliance with security requirements.
đşđ¸ United States â Remote
đľ $110k - $130k / year
â° Full Time
đ Senior
đ´ Lead
đˇ Infrastructure Engineer
đ April 28
Infrastructure Engineer building infrastructure and developing CI/CD for clinical AI/ML platform at Bayesian Health. Collaborating with cross-functional teams to ensure reliable and scalable systems.
đ April 25
Senior Cloud Infrastructure Engineer for HealthMark Group focusing on building and managing cloud infrastructure solutions. Ensure performance, uptime, and security for cloud native solutions.
đşđ¸ United States â Remote
đľ $110k - $140k / year
â° Full Time
đ Senior
đˇ Infrastructure Engineer
đ April 21
Cloud Infrastructure Engineer responsible for designing and managing GCP infrastructure. Ensuring scalability, reliability, and security of systems in a federal context.
đşđ¸ United States â Remote
đ° $550k Series B - GamePlan Technologies on 2013-10
â° Full Time
đĄ Mid-level
đ Senior
đˇ Infrastructure Engineer
đ April 20
Senior Software Engineer managing GPU and cloud infrastructure for AI Humans at Tavus. Collaborating across teams for reliable deployments and empowering engineering experiences.