Search Remote Jobs

Senior Solutions Architect, Cloud Partner Operations

šŸ”„ 0 minutes ago

šŸ„ California – Remote

infoinfo

šŸ’µ $224k - $356.5k / year

ā° Full Time

🟠 Senior

šŸ’» Solutions Engineer

šŸ¦… H1B Visa Sponsor

infoinfo

šŸ‘» Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

šŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

šŸ„ Healthcare

šŸ­ Manufacturing

šŸ¤– Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

šŸ“‹ Description

• Solve hard Day 2 operations problems at scale alongside partner engineers • Find causes, prototype approaches, validate solutions under representative load, and leave operational practices partners can run • Help partners prepare operating models for new NVIDIA platforms, capacity, services, and use cases • Drive adoption in live environments without degrading service • Improve reliability, performance, and economics using incident frequency, recovery time, utilization, and cost-per-token measures • Identify and help close Day 2 maturity gaps across people, process, tooling, telemetry, security, and incident response • Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows • Spot cross-partner patterns and provide field evidence to account teams, support, product, and engineering • Improve NVIDIA's factory planning function

šŸŽÆ Requirements

• BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure • Experience building, operating, or improving distributed infrastructure under real production load • Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure • Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling • Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation • Detailed evidence-led troubleshooting across system boundaries • Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority • Strong communication, prioritization, and time-management skills across multiple partner engagements • Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load • Experience building or maturing a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design • Hands-on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72, or NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and GPU or Network Operators • Experience improving fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation

šŸ–ļø Benefits

• Competitive salaries • Generous benefits package • Equity • Benefits

Apply Now

Similar Jobs

šŸ”„ 43 minutes ago

Ascension

10,000+ employees

šŸ„ Healthcare

šŸ¤ Non-profit

Enterprise Solution Architect modernizing healthcare IT systems for Ascension, a nonprofit Catholic health system. Governing architecture, integration, data, cloud, security, and AI initiatives.

šŸ”„ 46 minutes ago

Virtualitics

51 - 200

šŸŽ–ļø Defense

šŸ¤– Artificial Intelligence

ā˜ļø SaaS

Senior Solutions Architect guiding defense and government customers using Virtualitics’ AI-native readiness applications. Designing architectures, supporting pre-sales, and managing technical risks.

šŸ‡ŗšŸ‡ø United States – Remote

šŸ”„ Funding within the last year

šŸ’° $15M Debt Financing - Virtualitics on 2025-09

ā° Full Time

🟠 Senior

šŸ’» Solutions Engineer

šŸ”„ 49 minutes ago

Stand Together

5001 - 10000

šŸ„ Healthcare

šŸ’¼ Consulting

āš–ļø Legal

IAM Solutions Architect designing Okta and Auth0 identity security for Stand Together, a philanthropic community tackling America’s biggest social problems. Building human and non-human access governance and zero-trust systems.

šŸ”„ 50 minutes ago

Mercury Insurance

5001 - 10000

🚘 Automotive

šŸ’¼ Consulting

šŸ“¦ Logistics

Senior actuarial solutions engineer modernizing Mercury Insurance’s P&C actuarial applications and dashboards. Improving data access, pricing analysis, monitoring, reporting, and decision support.

šŸ”„ 1 hour ago

Traversal

11 - 50

šŸ¤– Artificial Intelligence

ā˜ļø SaaS

šŸ¤ B2B

Solutions Engineer guiding enterprise customers from technical discovery through deployment. Helping Traversal scale sales of its AI-powered SRE infrastructure platform.

šŸ‡ŗšŸ‡ø United States – Remote

šŸ’µ $150k - $300k / year

šŸ’° Seed on 2025-07

ā° Full Time

🟔 Mid-level

🟠 Senior

šŸ’» Solutions Engineer