
10,000+ employees
Founded 1993
đĽ Healthcare
đ Manufacturing
đ¤ Artificial Intelligence
Healthcare ⢠Manufacturing ⢠Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
đĽ 0 minutes ago
đ California, Washington â Remote
đľ $184k - $356.5k / year
â° Full Time
đ Senior
đ§âđť Full-stack Engineer
đŚ H1B Visa Sponsor
đť Ghost score 1%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
đĽ Healthcare
đ Manufacturing
đ¤ Artificial Intelligence
Healthcare ⢠Manufacturing ⢠Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
⢠Lead NVIDIA Cloud Partner Day 2 operational readiness efforts following initial deployment and activation ⢠Collaborate with NVIDIA Cloud Partners to establish systems, procedures, automation, and operational methods for accelerated infrastructure ⢠Build continuous validation for GPU, CPU, storage, and network health across large-scale AI clusters ⢠Establish telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, networking, storage, Kubernetes, and AI workloads ⢠Develop automated workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service ⢠Manage GPU fleet lifecycle, including drivers, firmware, Kubernetes nodes, OS patching, configuration management, upgrades, and configuration drift ⢠Translate NVIDIA NCP requirements and reference architectures into production operating practices, validation criteria, runbooks, automation, and measurable standards ⢠Define health signals, SLOs, metrics, acceptance criteria, and infrastructure readiness validation ⢠Build reusable tooling, automation, implementation guides, runbooks, playbooks, and reference implementations across NCP environments
⢠BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience ⢠8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments ⢠Strong experience operating Linux-based distributed systems and cloud infrastructure in production ⢠Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments ⢠Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements ⢠Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management ⢠Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers ⢠Programming and automation experience using Python, Go, shell scripting, or similar languages ⢠Experience managing extensive GPU or accelerated computing infrastructure that supports AI training and inference workloads ⢠Experience with NVIDIA technologies including DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software ⢠Proven experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure and operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency ⢠Extensive knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines ⢠Knowledge of failure modes related to large distributed AI workloads and the infrastructure features necessary to consistently support extended training and production inference
⢠Equity ⢠Benefits
Apply Nowđ Yesterday
Senior backend engineer building scalable PHP Laravel services for ShippyProâs shipping and fulfillment platform. Architecting microservices, distributed workflows, and AI automation for global merchants.
đşđ¸ United States â Remote
đľ âŹ42k - âŹ56k / year
đ° $15M Series B - ShippyPro on 2023-11
â° Full Time
đ Senior
đ§âđť Full-stack Engineer
đ Yesterday
Full Stack Engineer building React/Node cloud applications for ICFâs healthcare quality reporting systems. Developing microservices, serverless solutions, and analytics supporting CMS.
đşđ¸ United States â Remote
đľ $98.6k - $167.6k / year
đ° $29M Grant on 2023-03
â° Full Time
đĄ Mid-level
đ Senior
đ§âđť Full-stack Engineer
đ Yesterday
Cisco technical leader shaping C++/Python testing infrastructure, CI/CD, and developer tooling. Driving quality standards and engineering productivity across teams.
đşđ¸ United States â Remote
đľ $177.6k - $257.4k / year
â° Full Time
đ Senior
đ§âđť Full-stack Engineer
đŚ H1B Visa Sponsor
đ Yesterday
Cisco technical leader shaping C++/Python testing ecosystems, CI/CD quality gates, and developer platforms. Mentoring engineers and driving testing standards across Cisco.
đşđ¸ United States â Remote
đľ $177.6k - $257.4k / year
â° Full Time
đ Senior
đ§âđť Full-stack Engineer
đŚ H1B Visa Sponsor
đ Yesterday
Senior software developer building secure, scalable applications for OnTrac, a U.S. same-day and next-day delivery provider. Developing, deploying, troubleshooting, and improving software solutions.
đşđ¸ United States â Remote
đľ $104.8k - $131k / year
â° Full Time
đ Senior
đ§âđť Full-stack Engineer