Senior HPC DevOps Engineer

🕒 July 13

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, implement, and maintain large-scale HPC/AI clusters with state-of-the-art monitoring, logging, and alerting systems. • Utilize and develop tools to manage infrastructure as code, ensuring scalable and repeatable deployments. • Develop and maintain continuous integration and continuous delivery (CI/CD) pipelines to automate deployment processes. • Develop automation scripts and tools to automate deployment, configuration management, and operational monitoring. • Perform comprehensive troubleshooting from bare metal to application level, ensuring system reliability and efficiency. • Serve as a technical resource, developing and sharing best practices with internal teams. • Support R&D activities and engage in proof of concepts (POCs) and proof of values (POVs) for future improvements.

🎯 Requirements

• B.Sc. in Computer Science, Engineering, or a related field with 5+ years of experience. • Deep knowledge of HPC and AI solution technologies, including CPUs, GPUs, high-speed interconnects, and supporting software. • Advanced proficiency in programming and scripting languages, with a solid understanding of object-oriented programming principles. • Familiarity with Jenkins, Ansible, Puppet/Chef. • Excellent knowledge of Windows and Linux (Redhat/CentOS and Ubuntu), networking and OS-level security. • Deep understanding of networking protocols such as InfiniBand and Ethernet. • Experience with job scheduling workloads and orchestration tools such as Slurm and Kubernetes. • Background with multiple storage solutions like Lustre, GPFS, ZFS, and XFS. • Expertise with virtual systems (VMware, Hyper-V, KVM, Citrix). • Familiarity with cloud platforms (AWS, Azure, Google Cloud).

🏖️ Benefits

• Health insurance • Professional development opportunities • Flexible work arrangements

Apply Now

Similar Jobs

🕒 July 13

ClickHouse

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Senior Site Reliability Engineer at ClickHouse responsible for maintaining reliability and performance of cloud infrastructure. Collaborating with engineering teams to design scalable systems for real-time analytics.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Puppet

Python

SQL

Terraform

Go

🕒 July 13

JUPUS

51 - 200

⚖️ Legal

💼 Consulting

🤖 Artificial Intelligence

DevOps Lead overseeing infrastructure strategy for a rapidly scaling AI legal tech platform. Collaborating with engineering teams to modernize infrastructure and drive best practices.

AWS

Cloud

Grafana

Kubernetes

Terraform

🕒 July 11

Dedalus

5001 - 10000

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

DevOps Automation Specialist developing automation logic with VMware vRealize Automation. Working remotely for Dedalus, a global leader in healthcare technology.

🗣️🇩🇪 German Required

Ansible

Linux

Oracle

Postgres

Puppet

Python

SaltStack

VMware

🕒 July 9

XTEL

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Join XTEL as a Senior DevOps Engineer to build Azure infrastructure and automate deployments. Collaborate across teams in a remote, inclusive, growth-focused environment.

Azure

Cloud

DNS

Kubernetes

TCP/IP

Terraform

🕒 July 7

Digistore24 USA

51 - 200

📣 Marketing

💼 Consulting

☁️ SaaS

DevOps Generalist automating infrastructure tasks for a fast-growing tech company. Collaborate with international teams to enhance system reliability and performance.

🗣️🇩🇪 German Required

Cloud

Kubernetes