Research Scientist / Engineer – Reinforcement Learning Infrastructure

🔥 12 hours ago

🇪🇺 Europe – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Luma AI

Luma AI

2 - 10 employees

🤖 Artificial Intelligence

☁️ SaaS

📱 Media

💰 $90M Series B on 2025-01

Artificial Intelligence • SaaS • Media

Luma AI is a Palo Alto–based company building multimodal generative AI models and creative workflow tools focused on video, image, audio, and text production. Luma offers models and products such as Ray (video generation), UNI-1/Uni-1. 1 (brand intelligence models), and an API/enterprise offerings that allow creative teams to generate, edit, and direct cinematic-quality assets with brand continuity. They emphasize frontier research (Open Physical AI Lab), in-house model development, and products for creative teams, agencies, and production services. Luma also provides APIs, enterprise plans, a creative partner program, and learning resources.

📋 Description

• Design, build, and scale distributed RL post-training systems across thousands of GPUs • Build high-throughput rollout generation integrating vLLM, SGLang, weight synchronization, and asynchronous/off-policy schemes • Design scalable RL environments for agentic, multi-step tasks including sandboxed code execution, tool use, computer use, and multimodal interaction • Build reward infrastructure including verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking • Develop evaluation, monitoring, and debugging tooling for stable large-scale RL runs • Advance training efficiency and stability and turn post-training ideas into production runs with researchers • Learn the current RL stack, diagnose bottlenecks, ship and validate improvements, and harden the full loop across thousands of GPUs

🎯 Requirements

• Hands-on experience post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale • Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use • Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang) • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads • Containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads • Research contributions in RL for LLMs, or open-source contributions to RL training frameworks

🏖️ Benefits

• Equal opportunity employer • Flexible remote work arrangement in the EU

Apply Now

Similar Jobs

🕒 2 days ago

Synthesia

501 - 1000

📣 Marketing

💼 Consulting

📦 Logistics

Research Scientist developing diffusion models for natural, real-time interactive avatars. Synthesia builds AI video technology for business communication and enterprise skill development.

🇪🇺 Europe – Remote

🔥 Funding within the last year

💰 $200M Series E - Synthesia on 2025-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist