Senior Systems Engineer, Storage – DGX Cloud

🕒 June 8

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🤖 Artificial Intelligence

🎮 Gaming

Artificial Intelligence • Gaming • Automotive

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, deploy, and operate solutions on Kubernetes for large-scale storage and data platforms, including the manifests, Helm charts, and operators that run them. • Build tools, services, and automation that improve the lifecycle of storage and data systems – from provisioning and configuration through deployment, scaling, and day-2 operations. • Develop and operate telemetry and observability for production systems – metrics, logging, tracing, dashboards, and alerting – so that system health, availability, and latency are measurable and actionable. • Apply strong analytical troubleshooting skills to diagnose and resolve complex issues across distributed, containerized infrastructure. • Work closely with peers and partner teams to improve the lifecycle of services, from inception and design through deployment, operation, and refinement. • Scale systems sustainably through automation, infrastructure-as-code, and CI/CD, and evolve systems by pushing for changes that improve reliability and velocity. • Support services before they go live through activities such as deployment automation, capacity planning, and launch and readiness reviews. • Practice sustainable incident response and postmortems, and participate in an on-call rotation to support production systems.

🎯 Requirements

• BS degree (or equivalent experience) in Computer Science or related technical field involving coding. • 12+ years of practical experience. • Hands-on experience with Kubernetes – deploying, configuring, and operating workloads and solutions on Kubernetes in production. • Experience building tools and services for storage, data, or platform infrastructure, with solid software design fundamentals (algorithms, data structures, complexity analysis) on large-scale Linux-based systems. • Experience building and operating telemetry and observability using tools such as Prometheus, InfluxDB, Grafana, and the Elastic stack. • Strong analytical troubleshooting skills with a systematic, root-cause-driven approach to identifying and resolving complex problems. • Proficiency in one or more of the following: Python, Go, or Java. • Good knowledge of infrastructure configuration management and infrastructure-as-code tools such as Ansible, Chef, Puppet, ArgoCD, Git Pipelines, and Terraform.

🏖️ Benefits

• Equity • Health insurance • Retirement plans • Paid time off • Professional development opportunities

Apply Now

Similar Jobs

🕒 June 8

Datavant

201 - 500

⚕️ Healthcare Insurance

☁️ SaaS

🏢 Enterprise

Senior Systems Analyst supporting Oracle HCM technical initiatives at healthcare data collaboration platform Datavant. Focus on integrations, reporting, and system improvements with compliance adherence.

🕒 June 6

Amgen

10,000+ employees

🧬 Biotechnology

💊 Pharmaceuticals

🔬 Science

Business Systems Analyst at Amgen focusing on analyzing requirements and designing efficient IT solutions. Collaborating with teams to ensure operational efficiency and successful delivery aligned to business goals.

🕒 June 5

Autodesk

10,000+ employees

📱 Media

Senior Search Systems Engineer at Autodesk transforming structured marketing data into AI workflows. Collaborating with data engineering and analytics teams to enhance SEO and AEO strategy.

🕒 June 5

AI & Data Systems Engineer responsible for implementing internal AI tools at Ceribell. Collaborating with stakeholders and managing infrastructure for data systems in healthcare technology.

🇺🇸 United States – Remote

💵 $141k - $173k / year

💰 $50M Series C on 2022-09

⏰ Full Time

🟡 Mid-level

🟠 Senior

⚙️ Systems Engineer

🕒 June 5

T-Rex Solutions, LLC

201 - 500

🔒 Cybersecurity

🏛️ Government

Senior Business Systems Analyst supporting the General Services Administration modernization effort. Engaging in product discovery and requirements gathering for cutting-edge technologies.