Search Remote Jobs

Senior HPC AI Cluster Engineer

🔥 0 minutes ago

🏄 California – Remote

infoinfo

💵 $176k - $333.5k / year

⏰ Full Time

🟠 Senior

🤖 Artificial Intelligence

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, implement, and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting • Manage Linux job/workload schedules and orchestration tools • Develop and maintain continuous integration and delivery pipelines • Develop tooling to automate deployment and management of large-scale infrastructure environments • Automate operational monitoring and alerting and enable self-service consumption of resources • Deploy monitoring solutions for servers, network, and storage • Troubleshoot from bare metal through the operating system, software stack, and application level • Develop, redefine, and document standard methodologies for internal teams • Support Research & Development activities and engage in POCs/POVs for future improvements • Interact with HPC, OS, GPU compute, and systems specialists to architect, develop, and bring up large-scale performance platforms • Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs • Work with accelerated computing and deep learning software and hardware platforms, researchers, developers, and customers to improve workflows and develop differentiated solutions

🎯 Requirements

• A degree in Computer Science, Engineering, or a related field (or equivalent experience) • 8+ years of experience • Knowledge of HPC and AI solution technologies from CPUs and GPUs to high-speed interconnects and supporting software • Experience with job scheduling workloads and orchestration tools such as Slurm and Kubernetes • Excellent knowledge of Windows and Linux (Red Hat/CentOS and Ubuntu) networking and internals, ACLs, OS-level security protection, and common protocols such as TCP, DHCP, and DNS • Experience with storage solutions such as Lustre, GPFS, and Weka.io • Familiarity with newer and emerging storage technologies • Python programming and Bash scripting experience • Experience with automation and configuration management tools such as Jenkins, Ansible, Puppet, and Chef • Deep knowledge of networking protocols such as InfiniBand and Ethernet • Deep understanding and experience with virtual systems such as VMware, Hyper-V, KVM, or Citrix • Familiarity with cloud computing platforms such as AWS, Azure, and Google Cloud • Knowledge of CPU and/or GPU architecture • Knowledge of Kubernetes and container-related microservice technologies • Experience with GPU-focused hardware/software such as DGX and CUDA • Experience with RDMA (InfiniBand or RoCE) fabrics

🏖️ Benefits

• Competitive salaries • Generous benefits package • Equity

Apply Now

Similar Jobs

🔥 1 hour ago

Databricks

1001 - 5000

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Senior technology partnerships leader connecting AI code-generation and app-building platforms to Databricks’ Data and AI Platform. Driving integrations, joint solutions, and partner-led growth.

🔥 2 hours ago

CVS Health

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🛒 Retail

Lead long-term conversational AI product strategy for Signify Health, a CVS Health healthcare services business. Shape outreach, rescheduling, and care-continuity solutions.

🔥 17 hours ago

Twilio

5001 - 10000

🔌 API

🤝 B2B

Senior Data and AI Specialist optimizing Twilio’s CPaaS telecommunications network. Building LLM agents, Python workflows, and data products to improve quality, reduce fraud, and optimize margins.

🕒 Yesterday

TEKsystems

10,000+ employees

💼 Consulting

🎯 Recruiter

🏢 Enterprise

Practice Architect II leading AI/ML consulting and delivery programs for TEKsystems’ enterprise technology-services clients. Managing solution design, profitability, teams, risks, and customer outcomes.

🕒 Yesterday

Connection

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

AI Power Platform Administrator governing Microsoft Power Platform, Copilot, security, and automation solutions. Supporting Connection’s IT hardware, software, cloud, and technology services business.