Search Remote Jobs

Senior Engineer, NCX

🕒 September 1

🌐 Germany, Spain, +1 more countries – Remote

infoinfo

💵 zł292.5k - zł650k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Lead NVIDIA Cloud Partner Day 2 operational readiness efforts. • Collaborate with NVIDIA Cloud Partners to establish systems, procedures, automation, and operational methods for managing NVIDIA accelerated infrastructure after deployment and activation. • Develop continuous validation methods for GPU, CPU, storage, and network health across large-scale AI clusters. • Establish telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. • Build automated workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service. • Implement scalable GPU fleet lifecycle strategies, including driver and firmware administration, Kubernetes node maintenance, OS patching, configuration management, upgrades, and configuration drift identification. • Translate NVIDIA NCP requirements and reference architectures into production practices, validation criteria, runbooks, automation, and measurable operational standards. • Develop health signals, SLOs, metrics, acceptance criteria, and ongoing validation mechanisms for infrastructure reliability and service readiness. • Develop reusable tooling, automation, implementation guides, runbooks, operational playbooks, and reference implementations across multiple NCP environments.

🎯 Requirements

• BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience. • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments. • Strong experience operating Linux-based distributed systems and cloud infrastructure in production. • Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments. • Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements. • Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management. • Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers. • Programming and automation experience using Python, Go, shell scripting, or similar languages. • Experience managing extensive GPU or accelerated computing infrastructure that supports AI training and inference workloads. • Experience with NVIDIA technologies including DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software. • Proven experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure and operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency. • Extensive knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines. • Knowledge of failure modes related to large distributed AI workloads and the infrastructure features necessary to consistently support extended training and production inference.

Apply Now

Similar Jobs

🕒 August 30

RockstarDevelopers GmbH

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior Fullstack Entwickler fĂźr Java- und Angular-Projekte bei Rockstardevelopers, einem deutschen Softwareentwickler. Entwicklung in Enterprise- und regulierten Kundenumgebungen mit modernen AI-Tools.

🇩🇪 Germany – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🗣️🇩🇪 German Required

🕒 August 30

Hypoport SE

1001 - 5000

🛡️ Insurance

🏥 Healthcare

💼 Consulting

Fullstack Developer building Angular, Java, APIs, and frontend solutions for Europace, Germany’s largest real-estate finance marketplace. Operating end-to-end services for consumer lending products.

🇩🇪 Germany – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🗣️🇩🇪 German Required

🕒 August 30

Hypoport SE

1001 - 5000

🛡️ Insurance

🏥 Healthcare

💼 Consulting

Tech Lead fĂźr Java/Spring-Boot-Ratenkreditplattform bei Europace, Deutschlands FinTech-Transaktionsplattform. Verantwortung fĂźr Architektur, Delivery, Betrieb und technische FĂźhrung.

🇩🇪 Germany – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🗣️🇩🇪 German Required

🕒 August 29

CPU Softwarehouse AG

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Full-Stack-Entwickler fĂźr Bankensoftware bei CPU Consulting & Software. Umsetzung moderner Frontend- und Backend-LĂśsungen mit Angular, TypeScript, Java und Spring Boot.

🇩🇪 Germany – Remote

💵 €60k - €95k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🗣️🇩🇪 German Required

🕒 August 28

HolzLand Becker GmbH

51 - 200

🛒 Retail

🏗️ Construction

Software Developer customizing Dynamics 365 Business Central for a home-and-garden eCommerce company. Managing AL development, APIs, SQL Server, Azure DevOps, and integrations.

🇩🇪 Germany – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🗣️🇩🇪 German Required

Azure

Cloud

MS SQL Server

SQL