Senior Solutions Architect, Cloud Partner Operations

🕒 August 21

🏄 California – Remote

infoinfo

💵 $224k - $356.5k / year

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Solve hard Day 2 operations problems at scale alongside partner engineers • Find causes, prototype approaches, validate solutions under representative load, and leave operational practices partners can run • Help partners prepare operating models for new NVIDIA platforms, capacity, services, and use cases • Drive adoption in live environments without degrading service • Improve reliability, performance, and economics using incident frequency, recovery time, utilization, and cost-per-token measures • Identify and help close Day 2 maturity gaps across people, process, tooling, telemetry, security, and incident response • Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows • Spot cross-partner patterns and provide field evidence to account teams, support, product, and engineering • Improve NVIDIA's factory planning function

🎯 Requirements

• BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure • Experience building, operating, or improving distributed infrastructure under real production load • Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure • Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling • Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation • Detailed evidence-led troubleshooting across system boundaries • Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority • Strong communication, prioritization, and time-management skills across multiple partner engagements • Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load • Experience building or maturing a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design • Hands-on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72, or NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and GPU or Network Operators • Experience improving fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation

🏖️ Benefits

• Competitive salaries • Generous benefits package • Equity • Benefits

Apply Now

Similar Jobs

🕒 August 21

Virtualitics

51 - 200

🎖️ Defense

🤖 Artificial Intelligence

☁️ SaaS

Senior Solutions Architect guiding defense and government customers using Virtualitics’ AI-native readiness applications. Designing architectures, supporting pre-sales, and managing technical risks.

🇺🇸 United States – Remote

🔥 Funding within the last year

💰 $15M Debt Financing - Virtualitics on 2025-09

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

🕒 August 21

Stand Together

5001 - 10000

🏥 Healthcare

💼 Consulting

⚖️ Legal

IAM Solutions Architect designing Okta and Auth0 identity security for Stand Together, a philanthropic community tackling America’s biggest social problems. Building human and non-human access governance and zero-trust systems.

🇺🇸 United States – Remote

💵 $175k - $205k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

💻 Solutions Engineer

🦅 H1B Visa Sponsor

infoinfo

🕒 August 21

Traversal

11 - 50

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

Solutions Engineer guiding enterprise customers from technical discovery through deployment. Helping Traversal scale sales of its AI-powered SRE infrastructure platform.

🇺🇸 United States – Remote

💵 $150k - $300k / year

💰 Seed on 2025-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

💻 Solutions Engineer

🕒 August 21

Loancrate

11 - 50

💳 Fintech

🤖 Artificial Intelligence

🤝 B2B

Senior Software Engineer productizing mortgage-lender onboarding with configuration, provisioning, and migration automation. Building AI-native mortgage workflow tooling with TypeScript, AWS, and Terraform.

🇺🇸 United States – Remote

💵 $170k - $300k / year

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

🕒 August 21

CTERA

51 - 200

🔒 Cybersecurity

Solutions Engineer supporting CTERA’s enterprise data and cloud storage platform across the Western US. Providing technical sales support, solution design, presentations, proofs of concept, demos, and training.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

💻 Solutions Engineer