Senior Solutions Architect – Large Scale AI Inference

Job not on LinkedIn

🔥 12 hours ago

🌐 France, United Kingdom, +4 more countries – Remote

infoinfo

💵 zł292.5k - zł507k / year

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Guide EMEA AI Natives customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters • Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workload among thousands of GPUs • Improve inference efficiency across quantization, speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large MoE deployments • Collaborate with NVIDIA product teams, including Dynamo, TensorRT-LLM, and NIXL, to accelerate customer success • Animate the AI inference developer community across EMEA through technical workshops, hackathons, and reference architectures • Establish technical direction for scalable, high-performance AI inference across demanding production environments

🎯 Requirements

• MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience • 5+ years of experience in Neural Networks inference optimization • Solid understanding of transformers inference optimization, including quantization, disaggregated inference, speculative decoding, continuous batching, and KV cache optimization • Practical experience in MoE inference at scale, including expert parallelism, WideEP, all-to-all communication, routing overhead, and load balancing at scale • Ability to engage effectively with ML engineers, researchers, and systems architects at a deep technical level • Hands-on experience with NVIDIA Dynamo, NIXL, Grove, or emerging disaggregated inference tooling • Understanding of GPU memory hierarchies and high-speed interconnects, including NVLink, InfiniBand, RDMA, and UCX • Contributions to advanced AI labs or large-scale AI infrastructure providers performing inference on thousands of GPUs • Published work or benchmarks in large-scale AI inference

🏖️ Benefits

• Highly competitive salaries • Comprehensive benefits package

Apply Now

Similar Jobs

🕒 4 days ago

Capgemini

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Ingénieur intégration maintenant un ERP critique chez Capgemini, partenaire de la transformation business et technologique. Déploiement applicatif, incidents techniques et modernisation des plateformes.

🗣️🇫🇷 French Required

Docker

ERP

Linux

Oracle

SQL

Unix

🕒 August 27

AssessFirst 🦄

51 - 200

👥 HR Tech

🤖 Artificial Intelligence

☁️ SaaS

Solution Engineer turning customer data into churn scoring, dashboards, and automation for AssessFirst’s global HR Tech platform. Influencing product, marketing, and business strategy through actionable insights.

🗣️🇫🇷 French Required

Vue.js

🕒 August 21

Automat-it

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior Solutions Architect designing AWS solutions for Automat-it’s startup customers in Paris. Leading pre-sales architecture, cloud optimization, and customer-facing AWS engagements.

🗣️🇫🇷 French Required

AWS

Cloud

Kubernetes

🕒 August 17

EUROPEAN DYNAMICS

501 - 1000

💼 Consulting

📦 Logistics

📣 Marketing

Solution Architect designing cloud, data, and software architectures for European Dynamics. Developing architecture models and customized artefacts while advising on modelling-tool governance for major European clients.

Cloud

Kafka

Postgres

🕒 August 11

Veeam Software

1001 - 5000

💼 Consulting

📦 Logistics

☁️ SaaS

Solutions Engineer evangelizing Veeam Vault and advising customers on data resilience, cyber recovery, and public cloud architectures. Driving technical wins, architecture reviews, and strategic adoption for Veeam’s data protection platform.

🇫🇷 France – Remote

💰 $500M Private Equity Round on 2019-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

💻 Solutions Engineer

AWS

Azure

Cloud

Vault

VMware