
1001 - 5000 employees
đ€ Artificial Intelligence
đą Enterprise
âïž SaaS
Artificial Intelligence âą Enterprise âą SaaS
Nebius Group is building one of the worldâs leading AI infrastructure companies, focusing on providing the necessary compute, storage, and tools for developers in the AI space. Based in Europe and listed on Nasdaq, Nebius has a global presence with R&D centers across Europe, North America, and Israel. The company's primary offering is an AI-centric cloud platform designed for intensive AI workloads, complemented by various other businesses involved in generative AI development, edtech, and autonomous technology.
đ„ 17 hours ago
Improve your chances of getting an interview by checking your resume score before you apply.

1001 - 5000 employees
đ€ Artificial Intelligence
đą Enterprise
âïž SaaS
Artificial Intelligence âą Enterprise âą SaaS
Nebius Group is building one of the worldâs leading AI infrastructure companies, focusing on providing the necessary compute, storage, and tools for developers in the AI space. Based in Europe and listed on Nasdaq, Nebius has a global presence with R&D centers across Europe, North America, and Israel. The company's primary offering is an AI-centric cloud platform designed for intensive AI workloads, complemented by various other businesses involved in generative AI development, edtech, and autonomous technology.
âą Own optimization work for specific model families, customer endpoints, or serving backends âą Run engine comparisons and recommend practical serving configurations for specific workloads âą Debug model quality or performance regressions during production rollouts âą Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token âą Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems âą Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery âą Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving âą Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token âą Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers âą Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations
âą Strong Python and PyTorch engineering skills âą Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems âą Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems âą Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving âą Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs âą Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams âą Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques âą Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods âą Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration âą CUDA or Triton familiarity âą Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects
âą Competitive compensation âą Career growth and learning opportunities âą Flexibility and ownership âą Collaborative and innovative culture âą Opportunity to work on impactful AI projects âą International environment and talented teams
Apply Nowđ July 8
Fullstack Engineer with a focus on Generative AI, developing a proactive cyber platform. Collaborate with top-tier experts in a fully remote setup.