Search Remote Jobs

Senior Solutions Architect, Cloud Partner Operations

🕒 August 13

🏄 California – Remote

infoinfo

💵 $224k - $356.5k / year

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 5%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Solve hard Day 2 operations problems at scale alongside partner engineers. • Find causes, prototype approaches, validate them under representative load, and leave partners with operable practices. • Prepare partners for new NVIDIA platforms, capacity, services, and use cases. • Drive adoption in live environments without degrading service. • Improve reliability, performance, utilization, recovery time, and cost per token. • Identify and help close maturity gaps across people, process, tooling, telemetry, security, and incident response. • Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows. • Identify cross-partner patterns and provide field evidence to account teams, support, product, and engineering. • Improve NVIDIA Cloud Partner Day 2 operations and ecosystem capability.

🎯 Requirements

• BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience. • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure. • Experience building, operating, or improving distributed infrastructure under real production load. • Deep expertise in at least one part of the Day 2 stack, with hands-on large-scale GPU, HPC, or cloud infrastructure experience. • Experience with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms. • Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or similar tooling. • Strong Linux knowledge. • Experience with Python, Bash, or similar scripting for automation. • Evidence-led troubleshooting across system boundaries. • Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority. • Strong communication, prioritization, and time-management skills across multiple partner engagements. • Preferred/standout experience operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; building 24/7 operations; NVIDIA rack-scale platforms; NVIDIA operations technologies; fleet health or unit economics improvements.

🏖️ Benefits

• Competitive salaries • Generous benefits package • Equity

Apply Now

Similar Jobs

🕒 August 13

Databricks

1001 - 5000

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Solutions Architect helping large enterprises adopt Databricks’ Data and AI Platform. Defining account strategies, building proofs of concept, and driving ML and AI adoption.

🕒 August 13

LangChain

11 - 50

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Solutions Engineer helping LangChain customers evaluate, build, and deploy production AI agents in Texas. Partnering with sales and engineering teams on technical wins, POCs, and rollouts.

🕒 August 13

CM.com

1001 - 5000

💼 Consulting

📣 Marketing

📦 Logistics

AI Sales Solutions Engineer translating CM.com’s AI-powered customer engagement platform into business value for US customers. Leading technical discovery, demos, solution design, and partner enablement.

🕒 August 13

Domino Data Lab

201 - 500

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Senior Director leading pre-sales for Domino’s platform, which helps enterprises build and operate AI and data science solutions. Qualifying major deals, developing SE talent, and scaling the pre-sales organization.

🕒 August 13

EITACIES Inc.

51 - 200

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

EUC/VDI Solutions Architect building Omnissa Horizon virtual desktop environments and Windows images on AWS. Designing infrastructure, identity integration, desktop pools, and operational runbooks.