Senior Technical Program Manager, AI Infrastructure – Capacity Operations

Job not on LinkedIn

🔥 0 minutes ago

🏄 California – Remote

info

💵 $168k - $322k / year

⏰ Full Time

🟠 Senior

🔧 Technical Program Manager

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Own intake, triage, routing, and request-quality standards for accelerator-capacity requests across engineering and research teams • Run quarterly and annual demand forecasts, change control, review preparation, and follow-through • Prepare capacity-planning and allocation reviews using demand, supply, commitments, readiness, workload timing, and business priorities • Track capacity from request and forecast through delivery, readiness, assignment, and productive use • Coordinate infrastructure dependencies including access, storage, data movement, networking, readiness checks, migration timing, and provisioning tickets • Maintain dashboards and source-data quality, including freshness checks, reconciliation, missing-input ownership, and retirement of repeated manual reporting • Own operating cadences, agendas, action logs, dependency tracking, decision records, risk registers, blocking-issue paths, and closure • Prepare concise leadership reporting covering facts, risks, decisions, options, recommended paths, owners, and due dates • Drive tooling, automation, and decision-support initiatives with human approval gates, audit evidence, and safe operating controls • Influence research, platform engineering, infrastructure, data, finance, operations, and leadership collaborators

🎯 Requirements

• BS, MS, PhD in Electrical Engineering, Computer Science, Computer Engineering, or similar, or equivalent experience • 7+ years of technical program management or closely related experience in AI/ML platforms, distributed systems, cloud infrastructure, compute capacity, or technically demanding engineering environments • Experience owning recurring operational programs involving scarce-resource trade-offs, multiple cadences, executive clarity, and overlapping peak periods • Strong program mechanics including intake, forecasting, review preparation, dependency management, decision and risk records, action closure, documentation, and status communication • Technical proficiency sufficient to understand infrastructure constraints, inspect requirements and metrics, and work credibly with engineers and researchers • Data proficiency to evaluate source quality, interpret dashboards, reconcile conflicting views, define operational measures, and identify missing ownership • Track record transforming ambiguous cross-functional work into durable operating systems with clear owners, decisions, and critical issue paths • Excellent written and verbal communication • Ability to influence without authority across research, platform engineering, infrastructure, data, finance, operations, and leadership collaborators • Experience with GPU or accelerator-capacity planning, utilization programs, cluster operations, workload bring-up, or large-scale AI training and inference environments • Experience with demand and supply roadmaps, normalization across accelerator types, allocation reviews, quota or priority management, and capacity migrations • Familiarity with batch scheduling, cluster management, cloud, or data-center environments • Experience with observability or dashboard platforms, data-quality controls, and automated infrastructure or service-operations reporting • Experience driving API-enabled workflow automation or decision-support tools with human approval gates, audit evidence, and safe operating controls

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 16 hours ago

F5

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Senior technical program manager leading F5’s enterprise cybersecurity transformation. Driving cross-functional programs that secure and deliver applications across global digital environments.

🔥 20 hours ago

sFOX

51 - 200

₿ Crypto

💸 Finance

💳 Fintech

Technical Program Manager delivering secure, scalable fintech initiatives for sFOX, a prime dealer aggregating liquidity across exchanges and OTC desks. Partnering with engineering, product, compliance, risk, security, and operations teams to manage complex programs and releases.

🕒 Yesterday

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Engineering Program Manager driving Cisco Serviceability initiatives across CX and Engineering. Improving product support, case deflection, service-request resolution, and service costs through tools, features, and metrics.

🕒 Yesterday

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Engineering Program Manager driving Cisco Serviceability initiatives across CX and Engineering. Improving product support, case deflection, service resolution, and service costs through cross-functional programs.

🕒 3 days ago

Unity

5001 - 10000

🏭 Manufacturing

💼 Consulting

📣 Marketing

Senior Technical Program Manager coordinating Unity’s strategic platform projects. Delivering immersive gaming experiences across TVs and mobile devices with engineering and external partners.