Research Scientist – Performance Optimization

Job not on LinkedIn

🔥 1 minute ago

🇪🇺 Europe – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Luma AI

Luma AI

2 - 10 employees

🤖 Artificial Intelligence

☁️ SaaS

📱 Media

💰 $90M Series B on 2025-01

Artificial Intelligence • SaaS • Media

Luma AI is a Palo Alto–based company building multimodal generative AI models and creative workflow tools focused on video, image, audio, and text production. Luma offers models and products such as Ray (video generation), UNI-1/Uni-1. 1 (brand intelligence models), and an API/enterprise offerings that allow creative teams to generate, edit, and direct cinematic-quality assets with brand continuity. They emphasize frontier research (Open Physical AI Lab), in-house model development, and products for creative teams, agencies, and production services. Luma also provides APIs, enterprise plans, a creative partner program, and learning resources.

📋 Description

• Profile and optimize GPU, CPU, and accelerator code for maximum utilization and minimal latency • Write high-performance PyTorch, Triton, and CUDA, including custom operations when needed • Develop fused kernels and leverage tensor cores and modern hardware features across platforms • Optimize model architectures and implementations for distributed multi-node production deployment • Build performance monitoring and analysis tools and automation • Research and implement cutting-edge optimization techniques for transformer models • Profile current training and inference paths during the first 30 days • Ship and validate a kernel or architecture optimization during days 30–60 • Build monitoring and automation to prevent performance regressions during days 60–90

🎯 Requirements

• Expert-level Triton/CUDA programming and GPU optimization • Strong PyTorch skills, including kernel development and custom operations • Proficiency with profiling tools, including NVIDIA Nsight, torch profiler, and custom tooling • Deep understanding of transformer architectures and attention mechanisms • Experience with compilers and exporters such as torch.compile, TensorRT, ONNX, or XLA (nice to have) • Experience optimizing inference workloads for latency and throughput (nice to have) • Knowledge of Triton compiler and kernel fusion techniques (nice to have) • Knowledge of warp-level intrinsics and advanced CUDA optimization (nice to have)

🏖️ Benefits

• Equal opportunity employer • Voluntary diversity and inclusion survey participation; refusal will not affect the job application

Apply Now

Similar Jobs

🕒 2 days ago

Synthesia

501 - 1000

📣 Marketing

💼 Consulting

📦 Logistics

Research Scientist developing diffusion models for natural, real-time interactive avatars. Synthesia builds AI video technology for business communication and enterprise skill development.

🇪🇺 Europe – Remote

🔥 Funding within the last year

💰 $200M Series E - Synthesia on 2025-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist