Senior Engineer

🕒 6 days ago

🏄 California, Washington – Remote

infoinfo

💵 $184k - $356.5k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Lead NVIDIA Cloud Partner Day 2 operational readiness efforts following initial deployment and activation • Collaborate with NVIDIA Cloud Partners to establish systems, procedures, automation, and operational methods for accelerated infrastructure • Build continuous validation for GPU, CPU, storage, and network health across large-scale AI clusters • Establish telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, networking, storage, Kubernetes, and AI workloads • Develop automated workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service • Manage GPU fleet lifecycle, including drivers, firmware, Kubernetes nodes, OS patching, configuration management, upgrades, and configuration drift • Translate NVIDIA NCP requirements and reference architectures into production operating practices, validation criteria, runbooks, automation, and measurable standards • Define health signals, SLOs, metrics, acceptance criteria, and infrastructure readiness validation • Build reusable tooling, automation, implementation guides, runbooks, playbooks, and reference implementations across NCP environments

🎯 Requirements

• BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments • Strong experience operating Linux-based distributed systems and cloud infrastructure in production • Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments • Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements • Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management • Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers • Programming and automation experience using Python, Go, shell scripting, or similar languages • Experience managing extensive GPU or accelerated computing infrastructure that supports AI training and inference workloads • Experience with NVIDIA technologies including DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software • Proven experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure and operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency • Extensive knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines • Knowledge of failure modes related to large distributed AI workloads and the infrastructure features necessary to consistently support extended training and production inference

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🕒 August 22

ShippyPro

51 - 200

📦 Logistics

☁️ SaaS

🛍️ eCommerce

Senior backend engineer building scalable PHP Laravel services for ShippyPro’s shipping and fulfillment platform. Architecting microservices, distributed workflows, and AI automation for global merchants.

🇺🇸 United States – Remote

💵 €42k - €56k / year

💰 $15M Series B - ShippyPro on 2023-11

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🕒 August 22

Gartner

10,000+ employees

📣 Marketing

📦 Logistics

🏥 Healthcare

Senior Gartner analyst shaping AI strategy, SDLC adoption, and ROI measurement for software engineering leaders. Delivering research, client guidance, presentations, and sales support.

🇺🇸 United States – Remote

💵 $172k - $202.5k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

🕒 August 22

RR Digital

51 - 200

📣 Marketing

⚖️ Legal

🤝 B2B

Full Stack Developer building React, TypeScript, Python, and AWS document automation platforms for RR Digital, a legal marketing agency. Developing scalable production software for law-firm marketing operations.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🕒 August 22

MetalBear

11 - 50

☁️ SaaS

🤝 B2B

Engineering Team Lead leading cloud integrations for MetalBear’s Kubernetes development platform, mirrord. Building developer tooling for fast, production-like testing against live clusters.

🇺🇸 United States – Remote

🔥 Funding within the last year

💰 $12.5M Seed Round - MetalBear on 2025-09

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🕒 August 22

Sky Betting & Gaming

1001 - 5000

🎲 Gambling

🎮 Gaming

👥 B2C

Senior full-stack engineer building Go services and TypeScript/React experiences for Fanatics’ global sports, commerce, collectibles, and betting platform. Owning architecture, reliability, and technical mentorship.

🇺🇸 United States – Remote

💵 $121.6k - $200k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer