Senior Cloud Infrastructure, DevOps Solutions Architect

đŸ”„ 1 minute ago

🌐 France, United Kingdom, +2 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

đŸ’» Solutions Engineer

đŸ‘» Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

đŸ„ Healthcare

🏭 Manufacturing

đŸ€– Artificial Intelligence

Healthcare ‱ Manufacturing ‱ Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

‱ Own full-solution validation on partner software stacks, including cluster-wide stability testing, real training-workload acceptance, and multi-day, multi-rack burn-in ‱ Minimise time from cluster handover to first production workload by coordinating hardware bring-up, managed-service intake, and partner operations teams ‱ Own Day 2 production stability at fleet scale, including monitoring, logging, workload orchestration, fault detection and remediation, preventive maintenance, and firmware and field-notice rollout campaigns ‱ Assess customer environments and operate heterogeneous open platforms including Kubernetes, KubeVirt, Slurm, and GPU-aware schedulers ‱ Integrate enterprise-grade networking and storage and enable third-party ISV workloads ‱ Provide consultative guidance and hands-on troubleshooting across bare metal, operating systems, software stacks, container platforms, networking, and storage ‱ Support R&D, proofs of concept, and proofs of value validating new features, architectures, and upgrade approaches ‱ Act as technical leader for assigned accounts ‱ Run structured knowledge transfer and enablement ‱ Produce runbooks, onboarding materials, and best-practice guides for partner teams

🎯 Requirements

‱ BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience ‱ 8+ years in managing scalable cloud environments and automation engineering roles ‱ Proven understanding of networking fundamentals and data centre architectures ‱ Hands-on experience managing HPC/AI clusters and NVIDIA GPU-accelerated infrastructure, including deployment, driver and CUDA toolkit management, optimisation, workload profiling, and troubleshooting ‱ Extensive Kubernetes experience for container orchestration, resource scheduling, and scaling in GPU-accelerated and HPC environments ‱ Experience with scheduler internals, batch schedulers such as Slurm, and mixed bare-metal/virtualised multi-tenant estates such as KubeVirt ‱ Deep knowledge of Linux, including RedHat and Ubuntu, OS-level security, and protocols ‱ Experience with storage solutions such as Lustre, GPFS, ZFS, XFS, and Kubernetes storage technologies ‱ Proficiency in Python and Bash scripting ‱ Experience with configuration management and Infrastructure-as-Code tools such as Ansible and Terraform ‱ Experience with GitOps-based cluster lifecycle and upgrade management for large fleets ‱ Experience with observability stacks such as Grafana, Loki, and Prometheus ‱ Ability to measure and improve MTBI and job goodput on large GPU clusters ‱ Experience with fault detection, drain and remediation workflows, SLO/error-budget definition, and post-incident review ‱ Strong consultative background leading architectural reviews and presenting to executive stakeholders ‱ Knowledge of CI/CD pipelines and container-based microservices architectures ‱ Experience with NVIDIA GPU and Network Operators and NVIDIA Base Command Manager ‱ Familiarity with DCGM, XID diagnostics, node-level health agents, and fleet-wide reliability intelligence ‱ Expertise in AI-native scheduling and inference frameworks on Kubernetes, such as KAI, Grove, Dynamo, and NVIDIA Cloud Functions ‱ Background with RDMA-based fabrics such as InfiniBand or RoCE ‱ Exposure to Cumulus Linux, SONiC, Spectrum-X fabrics, DPU/DOCA infrastructure services, and NVLink/NVSwitch partition operations is a strong plus

đŸ–ïž Benefits

‱ NVIDIA products and work on next-generation AI/HPC systems ‱ Exposure to large-scale infrastructure projects and advanced GPU/HPC technologies ‱ Customer, partner, and cross-functional collaboration ‱ Technical leadership, knowledge transfer, and enablement opportunities

Apply Now

Similar Jobs

🕒 September 4

Capgemini

10,000+ employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Ingénieur intégration maintenant un ERP critique chez Capgemini, partenaire de la transformation business et technologique. Déploiement applicatif, incidents techniques et modernisation des plateformes.

đŸ—ŁïžđŸ‡«đŸ‡· French Required

Docker

ERP

Linux

Oracle

SQL

Unix

🕒 August 27

AssessFirst 🩄

51 - 200

đŸ‘„ HR Tech

đŸ€– Artificial Intelligence

☁ SaaS

Solution Engineer turning customer data into churn scoring, dashboards, and automation for AssessFirst’s global HR Tech platform. Influencing product, marketing, and business strategy through actionable insights.

đŸ—ŁïžđŸ‡«đŸ‡· French Required

Vue.js

🕒 August 21

Automat-it

51 - 200

đŸ’Œ Consulting

📩 Logistics

📣 Marketing

Senior Solutions Architect designing AWS solutions for Automat-it’s startup customers in Paris. Leading pre-sales architecture, cloud optimization, and customer-facing AWS engagements.

đŸ—ŁïžđŸ‡«đŸ‡· French Required

AWS

Cloud

Kubernetes

🕒 August 17

EUROPEAN DYNAMICS

501 - 1000

đŸ’Œ Consulting

📩 Logistics

📣 Marketing

Solution Architect designing cloud, data, and software architectures for European Dynamics. Developing architecture models and customized artefacts while advising on modelling-tool governance for major European clients.

Cloud

Kafka

Postgres

🕒 August 11

Cribl

501 - 1000

☁ SaaS

Partner Solutions Engineer enabling Cribl’s Southern European cloud and SaaS partners. Delivering technical presales, architecture, training, and proof-of-value engagements.

đŸ—ŁïžđŸ‡«đŸ‡· French Required

đŸ—ŁïžđŸ‡Ș🇾 Spanish Required

Cloud

JavaScript

Linux