Senior Solutions Architect, Generative AI

🕒 Yesterday

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Collaborating closely with customers to maximize GPU utilization and end-to-end workload throughput while improving infrastructure reliability and reducing infrastructure costs. • Designing and optimizing large-scale AI clusters across GPU compute, high-performance networking, storage, workload scheduling, orchestration, and observability. • Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack. • Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch. • Leading proof-of-concepts and performance studies for large-scale AI infrastructure, developing benchmarking tools, automation, runbooks, and technical collateral as needed. • Partnering with NVIDIA’s engineering, product, and sales teams to secure design wins and drive innovative solutions based on customer requirements and field feedback.

🎯 Requirements

• BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another Engineering field, or equivalent experience. • 6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role. • Deep understanding of Linux systems, distributed computing, GPU architectures, and the hardware and software components of large-scale AI clusters. • Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using technologies such as InfiniBand, RoCE, or GPUDirect RDMA. • Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues. • Experience profiling AI workloads and identifying performance bottlenecks across compute, networking, storage, and orchestration layers. • Experience with cluster schedulers and orchestration platforms such as Kubernetes and Slurm, along with containers and production monitoring systems. • Proficiency with Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.

🏖️ Benefits

• equity • generous benefits package

Apply Now

Similar Jobs

🕒 Yesterday

Stitch Fix

5001 - 10000

📣 Marketing

📦 Logistics

💼 Consulting

Lead Integration Engineer integrating global HCM solutions and building scalable HR tech at Stitch Fix. Collaborating with cross-functional teams to refine business processes and enhance technology systems.

Cloud

SOAP

Tableau

🕒 Yesterday

Ntara

11 - 50

🏭 Manufacturing

📦 Logistics

📣 Marketing

Solutions Engineer at Ntara delivering technical solutions around PIM and digital asset management for clients. Balancing project work and client engagements within a remote environment.

ASP.NET

Java

JavaScript

VBA

.NET

🕒 Yesterday

Impact Advisors

501 - 1000

💼 Consulting

🏥 Healthcare

🤖 Artificial Intelligence

Enterprise Solution Architect at Impact Advisors developing system service offerings for healthcare consulting. Responsible for designing and deploying scalable applications while advising clients on technology solutions.

AWS

Azure

Cloud

DNS

EC2

ETL

Informatica

Terraform

🕒 Yesterday

Solera, Inc.

5001 - 10000

🚘 Automotive

💼 Consulting

📦 Logistics

Account Manager for dealer solutions at Solera, leading customer retention and product adoption in automotive SaaS. Collaborating with teams to enhance dealer success and operational improvements.

🕒 Yesterday

Clockwork Systems, Inc.

11 - 50

🎮 Gaming

Senior Solutions Engineer at Clockwork.io leveraging AI networking expertise to enhance customer engagement and deliver tailored solutions. Collaborating with sales and engineering teams for effective deployments.

🇺🇸 United States – Remote

💵 $170k - $250k / year

💰 $21M Series A on 2022-03

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

AWS

Azure

Cloud

Google Cloud Platform

Java

Kubernetes

Linux

Python

SQL

Go