
1001 - 5000 Mitarbeiter
🤖 Künstliche Intelligenz
🏢 Unternehmen
☁️ SaaS
Artificial Intelligence • Enterprise • SaaS
Die Nebius Group baut eines der weltweit führenden Unternehmen für KI-Infrastruktur auf und konzentriert sich darauf, die notwendige Rechenleistung, Speicherkapazität und Tools für Entwickler im KI-Bereich bereitzustellen. Mit Sitz in Europa und an der Nasdaq notiert verfügt Nebius über eine globale Präsenz mit F&E-Zentren in Europa, Nordamerika und Israel. Das zentrale Angebot des Unternehmens ist eine KI-zentrierte Cloud-Plattform, die für rechenintensive KI-Workloads ausgelegt ist, ergänzt durch verschiedene weitere Geschäftsbereiche in den Bereichen Generative KI, Edtech und autonome Technologien.
🕒 vor 1 Monat
🌐 Niederlande, Deutschland, +2 weitere Länder – Remote
⏰ Vollzeit
🟠 Senior
🤖 Machine-Learning-Entwickler
🗣️🇺🇸🇬🇧 Englisch erforderlich
Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

1001 - 5000 Mitarbeiter
🤖 Künstliche Intelligenz
🏢 Unternehmen
☁️ SaaS
Artificial Intelligence • Enterprise • SaaS
Die Nebius Group baut eines der weltweit führenden Unternehmen für KI-Infrastruktur auf und konzentriert sich darauf, die notwendige Rechenleistung, Speicherkapazität und Tools für Entwickler im KI-Bereich bereitzustellen. Mit Sitz in Europa und an der Nasdaq notiert verfügt Nebius über eine globale Präsenz mit F&E-Zentren in Europa, Nordamerika und Israel. Das zentrale Angebot des Unternehmens ist eine KI-zentrierte Cloud-Plattform, die für rechenintensive KI-Workloads ausgelegt ist, ergänzt durch verschiedene weitere Geschäftsbereiche in den Bereichen Generative KI, Edtech und autonome Technologien.
• Token Factory is a part of Nebius Cloud, one of the world's largest GPU clouds, running tens of thousands of GPUs. • We are building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. • Our mission is to maximize throughput, minimise latency, and optimise cost-per-token across tens of thousands of GPUs. • Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. • Squeezing the maximum performance for a wide range of LLM architectures at scale (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-5). • Inference engines support: Implement novel speculative decoding architectures, optimise components of various LLM designs (dense/MoE, autoregressive/parallel), and contribute to open-source inference engines. • Low Precision Training & Inference: Design and productionise low-precision (FP8, NVFP4/MXFP4) training and inference pipelines with measurable gains in throughput and cost-efficiency.
• A profound understanding of theoretical foundations of machine learning and transformer architecture. • Experience profiling GPU workloads using Nsight, PyTorch profiler, or similar tools • Understanding of GPU memory hierarchy and compute/memory tradeoffs • Familiarity with important ideas in LLM space, such as MHA, RoPE, KV-cache, Flash Attention, and quantisation • Understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.) • Strong software engineering skills (we mostly use Python) • Deep experience with modern deep learning frameworks • Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing • Strong communication and leadership abilities
• Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams
Jetzt Bewerben🕒 vor 4 Monaten
AI Machine Learning Specialist collaborating with top-tier clients in the petrochemical industry. Focusing on machine learning algorithms and project management in a remote role.
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 6 Monaten
Top-class ML Engineer at JetBrains handling AI and ML engineering stack. Designing ML/LLM solutions and collaborating with engineering and research teams.
🗣️🇺🇸🇬🇧 Englisch erforderlich
Kotlin
Python
🕒 vor 8 Monaten
ML Engineer designing and prototyping machine learning solutions to enhance developer tools at JetBrains Research. Collaborating on a variety of machine learning projects.
🗣️🇺🇸🇬🇧 Englisch erforderlich