Search Remote Jobs

Senior HPC AI Cluster Engineer

🕒 August 20

🏄 California – Remote

infoinfo

💵 $176k - $333.5k / year

⏰ Full Time

🟠 Senior

🤖 Artificial Intelligence

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 5%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, implement, and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting • Manage Linux job/workload schedules and orchestration tools • Develop and maintain continuous integration and delivery pipelines • Develop tooling to automate deployment and management of large-scale infrastructure environments • Automate operational monitoring and alerting and enable self-service consumption of resources • Deploy monitoring solutions for servers, network, and storage • Troubleshoot from bare metal through the operating system, software stack, and application level • Develop, redefine, and document standard methodologies for internal teams • Support Research & Development activities and engage in POCs/POVs for future improvements • Interact with HPC, OS, GPU compute, and systems specialists to architect, develop, and bring up large-scale performance platforms • Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs • Work with accelerated computing and deep learning software and hardware platforms, researchers, developers, and customers to improve workflows and develop differentiated solutions

🎯 Requirements

• A degree in Computer Science, Engineering, or a related field (or equivalent experience) • 8+ years of experience • Knowledge of HPC and AI solution technologies from CPUs and GPUs to high-speed interconnects and supporting software • Experience with job scheduling workloads and orchestration tools such as Slurm and Kubernetes • Excellent knowledge of Windows and Linux (Red Hat/CentOS and Ubuntu) networking and internals, ACLs, OS-level security protection, and common protocols such as TCP, DHCP, and DNS • Experience with storage solutions such as Lustre, GPFS, and Weka.io • Familiarity with newer and emerging storage technologies • Python programming and Bash scripting experience • Experience with automation and configuration management tools such as Jenkins, Ansible, Puppet, and Chef • Deep knowledge of networking protocols such as InfiniBand and Ethernet • Deep understanding and experience with virtual systems such as VMware, Hyper-V, KVM, or Citrix • Familiarity with cloud computing platforms such as AWS, Azure, and Google Cloud • Knowledge of CPU and/or GPU architecture • Knowledge of Kubernetes and container-related microservice technologies • Experience with GPU-focused hardware/software such as DGX and CUDA • Experience with RDMA (InfiniBand or RoCE) fabrics

🏖️ Benefits

• Competitive salaries • Generous benefits package • Equity

Apply Now

Similar Jobs

🕒 August 20

Databricks

1001 - 5000

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Senior technology partnerships leader connecting AI code-generation and app-building platforms to Databricks’ Data and AI Platform. Driving integrations, joint solutions, and partner-led growth.

🕒 August 19

Net at Work

201 - 500

💼 Consulting

🏭 Manufacturing

📦 Logistics

AI trainer helping Net at Work’s SMB clients adopt practical, responsible AI workflows. Delivering workshops, executive briefings, and hands-on enablement sessions.

🕒 August 18

ICF

5001 - 10000

💼 Consulting

🏛️ Government

🏥 Healthcare

Project Specialist supporting Alaska Native victim-services programs for ICF, a global advisory and technology services provider. Delivering culturally responsive training, compliance assistance, and strategic planning across Alaska.

🇺🇸 United States – Remote

💵 $61k - $103.6k / year

💰 $29M Grant on 2023-03

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

🕒 August 17

Xsolis

201 - 500

🏥 Healthcare

☁️ SaaS

🤖 Artificial Intelligence

AI Operations Manager ensuring reliable, observable, and secure traditional ML, GenAI, and agentic systems. Advancing ethical AI operations for Xsolis’s healthcare technology platform.

🕒 August 17

ICF

5001 - 10000

🏥 Healthcare

📦 Logistics

📣 Marketing

ICF project specialist supporting Alaska Native victim services programs through culturally responsive training, grant compliance, and strategic planning. Remote Alaska role serving communities and federal government clients.