Search Remote Jobs

Senior Solutions Architect, Cloud Partner Operations

šŸ”„ 0 minutes ago

šŸ„ California – Remote

info

šŸ’µ $224k - $356.5k / year

ā° Full Time

🟠 Senior

šŸ’» Solutions Engineer

šŸ¦… H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

šŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

šŸ„ Healthcare

šŸ­ Manufacturing

šŸ¤– Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

šŸ“‹ Description

• Solve hard Day 2 operations problems at scale alongside partner engineers. • Find causes, prototype approaches, validate them under representative load, and leave partners with operable practices. • Prepare partners for new NVIDIA platforms, capacity, services, and use cases. • Drive adoption in live environments without degrading service. • Improve reliability, performance, utilization, recovery time, and cost per token. • Identify and help close maturity gaps across people, process, tooling, telemetry, security, and incident response. • Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows. • Identify cross-partner patterns and provide field evidence to account teams, support, product, and engineering. • Improve NVIDIA Cloud Partner Day 2 operations and ecosystem capability.

šŸŽÆ Requirements

• BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience. • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure. • Experience building, operating, or improving distributed infrastructure under real production load. • Deep expertise in at least one part of the Day 2 stack, with hands-on large-scale GPU, HPC, or cloud infrastructure experience. • Experience with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms. • Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or similar tooling. • Strong Linux knowledge. • Experience with Python, Bash, or similar scripting for automation. • Evidence-led troubleshooting across system boundaries. • Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority. • Strong communication, prioritization, and time-management skills across multiple partner engagements. • Preferred/standout experience operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; building 24/7 operations; NVIDIA rack-scale platforms; NVIDIA operations technologies; fleet health or unit economics improvements.

šŸ–ļø Benefits

• Competitive salaries • Generous benefits package • Equity

Apply Now

Similar Jobs

šŸ”„ 31 minutes ago

phData

201 - 500

šŸ’¼ Consulting

šŸ„ Healthcare

šŸ­ Manufacturing

Senior Solutions Architect leading enterprise Data & AI architecture and delivery for phData, a remote-first consultancy. Advising clients, guiding engineers, and applying AI-first engineering practices.

šŸ”„ 32 minutes ago

Databricks

1001 - 5000

šŸ¤– Artificial Intelligence

šŸ¢ Enterprise

ā˜ļø SaaS

Solutions Architect helping large enterprises adopt Databricks’ Data and AI Platform. Defining account strategies, building proofs of concept, and driving ML and AI adoption.

šŸ”„ 59 minutes ago

Fortive

10,000+ employees

šŸ„ Healthcare

šŸ­ Manufacturing

šŸ“¦ Logistics

Senior Solutions Consultant implementing Gordian’s capital planning software for facilities clients. Leading requirements discovery, data migrations, training, project delivery, and customer success.

šŸ”„ 1 hour ago

Databricks

1001 - 5000

šŸ¤– Artificial Intelligence

šŸ¢ Enterprise

ā˜ļø SaaS

Senior Solutions Architect guiding enterprise retail, consumer goods, and travel customers adopting Databricks’ Data and AI Platform. Defining technical strategies, driving ML/AI adoption, and mentoring field engineering teams.

šŸ”„ 2 hours ago

LangChain

11 - 50

šŸ¤– Artificial Intelligence

šŸ¤ B2B

ā˜ļø SaaS

Solutions Engineer helping LangChain customers evaluate, build, and deploy production AI agents in Texas. Partnering with sales and engineering teams on technical wins, POCs, and rollouts.