Senior Solutions Architect – Large Scale Neural Networks Inference

Job not on LinkedIn

🔥 12 hours ago

🌐 France, United Kingdom, +3 more countries – Remote

infoinfo

💵 zł292.5k - zł650k / year

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Lead the inference strategy for a portfolio of EMEA AI Natives customers, guiding engagements from initial proof of concept to production-scale deployments • Identify inference challenges across customer deployments, including latency, efficiency, cost per token, memory utilization, and low-latency networking • Architect and optimize high-performance inference pipelines using NVIDIA Dynamo, TensorRT-LLM, vLLM, SGLang, and other inference backends • Improve GPU utilization and AI cluster efficiency • Translate customer insights and deployment patterns into actionable product feedback • Develop the roadmap for the NVIDIA stack, including Dynamo, TensorRT-LLM, and NIM • Define the technical direction for AI inference across EMEA • Align NVIDIA and customer organization stakeholders to influence strategic technology decisions for next-generation AI inference at scale

🎯 Requirements

• MS or PhD in Computer Science, Engineering, or equivalent experience in the field • 8+ years in AI/ML infrastructure, with deep expertise in LLM/VLM inference optimization and production deployment at scale • Deep understanding of transformer inference acceleration: quantization (INT4/FP8), speculative decoding, disaggregated inference, continuous batching, KV cache optimization, and WideEP for MoE models • Understanding of GPU memory hierarchies and low-latency networking and their influence on inference performance • Proven track record to lead technical initiatives • Excellent communication skills, effective with research scientists, infrastructure engineers, and executive team members • Experience with NVIDIA's inference stack, including TensorRT-LLM, Triton Inference Server, NIM, and NVIDIA Dynamo • Experience with GPU orchestration on Kubernetes • Experience operating inference at scale inside a frontier AI lab or hyperscale's inference team • Contributions to open-source inference projects such as vLLM, SGLang, KServe, or NVIDIA Dynamo

🏖️ Benefits

• Highly competitive salaries • Comprehensive benefits package

Apply Now

Similar Jobs

🕒 4 days ago

Capgemini

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Ingénieur intégration maintenant un ERP critique chez Capgemini, partenaire de la transformation business et technologique. Déploiement applicatif, incidents techniques et modernisation des plateformes.

🗣️🇫🇷 French Required

Docker

ERP

Linux

Oracle

SQL

Unix

🕒 August 27

AssessFirst 🦄

51 - 200

👥 HR Tech

🤖 Artificial Intelligence

☁️ SaaS

Solution Engineer turning customer data into churn scoring, dashboards, and automation for AssessFirst’s global HR Tech platform. Influencing product, marketing, and business strategy through actionable insights.

🗣️🇫🇷 French Required

Vue.js

🕒 August 21

Automat-it

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior Solutions Architect designing AWS solutions for Automat-it’s startup customers in Paris. Leading pre-sales architecture, cloud optimization, and customer-facing AWS engagements.

🗣️🇫🇷 French Required

AWS

Cloud

Kubernetes

🕒 August 17

EUROPEAN DYNAMICS

501 - 1000

💼 Consulting

📦 Logistics

📣 Marketing

Solution Architect designing cloud, data, and software architectures for European Dynamics. Developing architecture models and customized artefacts while advising on modelling-tool governance for major European clients.

Cloud

Kafka

Postgres

🕒 August 11

Veeam Software

1001 - 5000

💼 Consulting

📦 Logistics

☁️ SaaS

Solutions Engineer evangelizing Veeam Vault and advising customers on data resilience, cyber recovery, and public cloud architectures. Driving technical wins, architecture reviews, and strategic adoption for Veeam’s data protection platform.

🇫🇷 France – Remote

💰 $500M Private Equity Round on 2019-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

💻 Solutions Engineer

AWS

Azure

Cloud

Vault

VMware