Senior ML Engineer, Token Factory

Job not on LinkedIn

🕒 July 1

🌐 Germany, Israel, +2 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

👻 Ghost score 30%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Nebius Group

Nebius Group

1001 - 5000 employees

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Nebius Group is building one of the world’s leading AI infrastructure companies, focusing on providing the necessary compute, storage, and tools for developers in the AI space. Based in Europe and listed on Nasdaq, Nebius has a global presence with R&D centers across Europe, North America, and Israel. The company's primary offering is an AI-centric cloud platform designed for intensive AI workloads, complemented by various other businesses involved in generative AI development, edtech, and autonomous technology.

📋 Description

• Token Factory is a part of Nebius Cloud, one of the world’s largest GPU clouds, running tens of thousands of GPUs. • We are building an inference & fine-tuning platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to train & deploy at massive scale. • Enhancing fine-tuning methodologies - both LoRA-based and full-parameter - for cutting-edge LLMs (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-4.7), focusing on both model quality and training efficiency. • Identifying LLM inference bottlenecks to drive production speedups. • Building model training and evaluation pipelines in JAX for speculative decoding, experimenting with architectures (dense/MoE, auto-regressive/parallel), and deriving scaling laws to guide resource allocation. • Investigating low-precision (FP8, NVFP4/MXFP4) methodologies for supervised fine-tuning and reinforcement learning - spanning both inference and training - optimized for modern hardware

🎯 Requirements

• A profound understanding of theoretical foundations of machine learning and reinforcement learning. • Deep expertise in modern deep learning for language processing and generation • Experience with training large models on multiple computational nodes • Reasonable understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.) • Strong software engineering skills (we mostly use Python) • Deep experience with modern deep learning frameworks (we use JAX) • Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing • Strong communication and leadership abilities

🏖️ Benefits

• Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams

Apply Now

Similar Jobs

🕒 March 27

XO Life

11 - 50

💼 Consulting

📦 Logistics

📣 Marketing

AI/ML Engineer developing machine learning models and algorithms for a leading digital health platform. Collaborating with cross-functional teams to drive impactful solutions in healthcare.

Apache

AWS

Cloud

Docker

Google Cloud Platform

JavaScript

Keras

Kubernetes

Microservices

MongoDB

Numpy

Pandas

Python

PyTorch

Scikit-Learn

Spark

Tensorflow

TypeScript

🕒 March 26

super.AI

11 - 50

🤖 Artificial Intelligence

Machine Learning Engineer working on AI solutions for global enterprises, focusing on high-quality output and performance. Collaborating across teams to implement AI and machine learning technologies.

AWS

Azure

Cloud

Distributed Systems

Docker

Google Cloud Platform

Kubernetes

Python

🕒 March 24

Blackwall

51 - 200

🔒 Cybersecurity

🏢 Enterprise

🤝 B2B

Machine Learning Engineer developing and deploying AI-driven security features for web protection at Blackwall. Collaborating with cross-functional teams to enhance the efficiency of security systems.

🇩🇪 Germany – Remote

💰 $49M Series B - BlackWall on 2025-03

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

Python

PyTorch

Scikit-Learn

SQL

Tensorflow

🕒 March 24

Veeva Systems

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior Machine Learning Engineer developing scalable AI solutions for life sciences. Collaborating with cross-functional teams to enhance clinical data insights.

AWS

Cloud

Google Cloud Platform

Kubernetes

Python

Terraform

🕒 March 9

Janea Systems

11 - 50

🏢 Enterprise

Lead Machine Learning Engineer at Janea Systems designing scalable machine learning systems. Collaborating with global teams in a remote-first company to deliver AI solutions for enterprise clients.

AWS

Azure

Cloud

Distributed Systems

Docker

Google Cloud Platform

Kubernetes

Python

PyTorch

Tensorflow