
10,000+ employees
Founded 1993
🏥 Healthcare
🏠Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🔥 2 hours ago
🏄 California – Remote
đź’µ $144k - $230k / year
⏰ Full Time
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏠Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Build and deploy sophisticated AI-powered tools and products supporting operation and optimization of the global GeForce NOW service • Transform production data streams, including signals, metrics, and logs, into actionable intelligence • Automate root cause analysis for incidents and predict future service trends and patterns • Build and implement AI/ML tools to identify incident root causes and operational trends • Lead development of LLM- and agent-based systems to improve operational efficiency • Establish and maintain data management practices and pipelines for large-scale model-development data • Own and enhance LLM-based pipelines • Integrate knowledge of LLM progress into product development • Serve as an authority on AI frameworks and recommend platforms, toolsets, and architectures for long-term product sustainability
• B.S. in Computer Science, Statistics, or Engineering, or equivalent experience • 5+ years of experience • Strong proficiency in Python • Familiarity with Go or other systems languages is a plus • Practical experience building, optimizing, and deploying AI tools • Strong knowledge of AI developments and LLM-based platforms • Hands-on experience with Kubernetes and AWS cloud environments • Ability to distinguish meaningful AI advances from noise in technical decisions • Expertise in automation and large-scale data pipelines • Experience with monitoring and visualization tools such as Grafana • Excellent ability to handle, transform, and manage data sources and pipelines • Understanding of SRE principles and production-environment management is preferred • Experience with LLM improvement pipelines and recent LLM training developments is preferred • Knowledge of LLMs and AI models, including ability to recommend sustainable platforms and approaches
• Equity • Benefits • Competitive salary package • Inclusive and supportive work environment
Apply Now🔥 3 hours ago
Site Reliability Engineer operating AWS cloud infrastructure for SitusAMC’s real estate technology solutions. Improving reliability, automation, observability, security, and application migrations.
🇺🇸 United States – Remote
đź’µ $95k - $135k / year
đź’° Private equity on 2020-05
⏰ Full Time
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🔥 5 hours ago
Senior Reliability Engineer improving embedded protection, control, and software products for GE Vernova’s decarbonization mission. Driving testing, field analytics, grid reliability, and cybersecurity compliance.
🇺🇸 United States – Remote
đź’µ $162.9k - $244.3k / year
⏰ Full Time
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 5 hours ago
Senior Staff DevOps Engineer scaling AWS Kubernetes infrastructure for SailPoint’s identity security platform. Leading enterprise service mesh adoption, PCI-compliant operations, and cloud-native reliability across global teams.
🔥 6 hours ago
Senior DevOps Engineer building AWS/Kubernetes infrastructure and CI/CD systems for Pacvue’s commerce media platform. Improving reliability, security, observability, and developer productivity across engineering teams.
🔥 8 hours ago
Senior backend engineer building Affirm’s reliability platform for scalable buy-now-pay-later systems. Designing distributed services, AI-assisted tooling, and operational safeguards for production health.
🇺🇸 United States – Remote
đź’µ $173k - $233k / year
đź’° Post-IPO Equity on 2021-01
⏰ Full Time
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor