Senior Solutions Architect, Cloud Partner Operations

Job not on LinkedIn

🕒 2 days ago

🏄 California – Remote

info

💵 $224k - $356.5k / year

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Solve hard Day 2 operations problems at scale alongside partner engineers. • Find causes, prototype approaches, validate them under representative load, and leave partners with operable practices. • Prepare partners for new NVIDIA platforms, capacity, services, and use cases. • Drive adoption in live environments without degrading service. • Improve reliability, performance, utilization, recovery time, and cost per token. • Identify and help close maturity gaps across people, process, tooling, telemetry, security, and incident response. • Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows. • Identify cross-partner patterns and provide field evidence to account teams, support, product, and engineering. • Improve NVIDIA Cloud Partner Day 2 operations and ecosystem capability.

🎯 Requirements

• BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience. • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure. • Experience building, operating, or improving distributed infrastructure under real production load. • Deep expertise in at least one part of the Day 2 stack, with hands-on large-scale GPU, HPC, or cloud infrastructure experience. • Experience with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms. • Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or similar tooling. • Strong Linux knowledge. • Experience with Python, Bash, or similar scripting for automation. • Evidence-led troubleshooting across system boundaries. • Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority. • Strong communication, prioritization, and time-management skills across multiple partner engagements. • Preferred/standout experience operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; building 24/7 operations; NVIDIA rack-scale platforms; NVIDIA operations technologies; fleet health or unit economics improvements.

🏖️ Benefits

• Competitive salaries • Generous benefits package • Equity

Apply Now

Similar Jobs

🕒 2 days ago

Databricks

1001 - 5000

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Solutions Architect helping large enterprises adopt Databricks’ Data and AI Platform. Defining account strategies, building proofs of concept, and driving ML and AI adoption.

🇺🇸 United States – Remote

💵 $180k - $247.5k / year

💰 $1.6G Series H on 2021-08

⏰ Full Time

🟡 Mid-level

🟠 Senior

💻 Solutions Engineer

🦅 H1B Visa Sponsor

info

🕒 2 days ago

Fortive

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

📦 Logistics

Senior Solutions Consultant implementing Gordian’s capital planning software for facilities clients. Leading requirements discovery, data migrations, training, project delivery, and customer success.

🇺🇸 United States – Remote

💰 Post-IPO Equity on 2020-03

⏰ Full Time

🟡 Mid-level

🟠 Senior

💻 Solutions Engineer

🦅 H1B Visa Sponsor

info

SQL

SSIS

🕒 2 days ago

LangChain

11 - 50

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Solutions Engineer helping LangChain customers evaluate, build, and deploy production AI agents in Texas. Partnering with sales and engineering teams on technical wins, POCs, and rollouts.

🇺🇸 United States – Remote

💰 $25M Series A on 2024-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

💻 Solutions Engineer

🕒 2 days ago

Presidio

1001 - 5000

💼 Consulting

📦 Logistics

🤖 Artificial Intelligence

Senior AWS Solutions Architect leading pre-sales, migration, security, and IaC solutions for Presidio’s telecom, media, entertainment, gaming, and sports clients.

🇺🇸 United States – Remote

💰 Private equity on 2011-05

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

🦅 H1B Visa Sponsor

info

🕒 2 days ago

Presidio

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior AWS Solutions Architect leading pre-sales, migration, security, and IaC solutions for Presidio, a global cloud, AI, and cybersecurity services provider. Supporting enterprise modernization and proposal delivery.