Research Scientist / Engineer – Training Infrastructure

🔥 12 hours ago

🇪🇺 Europe – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Luma AI

Luma AI

2 - 10 employees

🤖 Artificial Intelligence

☁️ SaaS

📱 Media

💰 $90M Series B on 2025-01

Artificial Intelligence • SaaS • Media

Luma AI is a Palo Alto–based company building multimodal generative AI models and creative workflow tools focused on video, image, audio, and text production. Luma offers models and products such as Ray (video generation), UNI-1/Uni-1. 1 (brand intelligence models), and an API/enterprise offerings that allow creative teams to generate, edit, and direct cinematic-quality assets with brand continuity. They emphasize frontier research (Open Physical AI Lab), in-house model development, and products for creative teams, agencies, and production services. Luma also provides APIs, enterprise plans, a creative partner program, and learning resources.

📋 Description

• Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs • Research and implement advanced parallelization, including FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel • Build monitoring, visualization, and debugging tools for large-scale training runs • Optimize training stability, convergence, and resource utilization across massive clusters • Learn the current training stack and diagnose stability and utilization issues at scale • Deliver a parallelization or stability improvement that measurably benefits a real training run • Build monitoring and tooling to keep large runs reliable and efficient

🎯 Requirements

• Extensive distributed PyTorch training and parallelisms in foundation-model training • Deep understanding of GPU clusters, networking, and storage systems • Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization • Strong Linux systems administration and scripting (nice to have) • Experience managing training runs across 100+ GPUs (nice to have) • Experience with containerization, orchestration, and cloud infrastructure (nice to have) • Experience at the level of FSDP and multi-node training • Ability to work remotely in the EU

🏖️ Benefits

• Equal opportunity employer • Voluntary diversity and inclusion survey participation; refusal will not affect the job application • Remote work arrangement

Apply Now

Similar Jobs

🕒 2 days ago

Synthesia

501 - 1000

📣 Marketing

💼 Consulting

📦 Logistics

Research Scientist developing diffusion models for natural, real-time interactive avatars. Synthesia builds AI video technology for business communication and enterprise skill development.

🇪🇺 Europe – Remote

🔥 Funding within the last year

💰 $200M Series E - Synthesia on 2025-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist