
10,000+ employees
Founded 1993
đĽ Healthcare
đ Manufacturing
đ¤ Artificial Intelligence
Healthcare ⢠Manufacturing ⢠Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
đĽ 0 minutes ago
đ California â Remote
đľ $144k - $230k / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đť Ghost score 1%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
đĽ Healthcare
đ Manufacturing
đ¤ Artificial Intelligence
Healthcare ⢠Manufacturing ⢠Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
⢠Build and deploy sophisticated AI-powered tools and products supporting operation and optimization of the global GeForce NOW service ⢠Transform production data streams, including signals, metrics, and logs, into actionable intelligence ⢠Automate root cause analysis for incidents and predict future service trends and patterns ⢠Build and implement AI/ML tools to analyze production data, identify root causes for complex incidents, and identify future operational trends ⢠Lead development of LLM- and Agent-based systems to improve operational efficiency ⢠Establish and maintain data management practices and construct workflows for large-scale data sources vital for model development ⢠Take charge of and enhance LLM-based pipelines while integrating LLM progress into product development ⢠Act as an authority on AI frameworks and recommend platforms, toolsets, and architectural approaches for long-term technical sustainability
⢠B.S. in Computer Science, Statistics, or Engineering (or equivalent experience) ⢠5+ years of experience ⢠Strong proficiency in Python ⢠Familiarity with Go or other systems languages is a plus ⢠Practical experience building, optimizing, and deploying AI tools ⢠Strong knowledge of the AI space and current developments, including understanding how LLM-based platforms are built, optimized, and which platforms work best ⢠Hands-on experience with container orchestration (Kubernetes) and cloud environments (AWS cloud) ⢠Active engagement with developments in the AI field and ability to distinguish meaningful advances from noise when making technical decisions ⢠Expertise in automation and handling large-scale data pipelines ⢠Experience applying monitoring and visualization tools, such as Grafana, to interact with data ⢠Excellent ability to handle data sources and pipelines to transform and manage data ⢠Current experience in LLM improvement pipelines and a strong grasp of recent developments in LLM training ⢠Understanding of SRE concepts and experience managing production environments ⢠Experience with Kubernetes, AWS, and other cloud technologies ⢠Excellent knowledge of LLMs and AI Models ⢠Proficiency in automation
⢠Competitive salary package ⢠Equity ⢠Benefits
Apply NowđĽ 5 hours ago
Senior DevOps Engineer securing AWS/Azure cloud infrastructure for Koniag Government Services. Automating DevSecOps, CI/CD security, compliance, and incident response for federal customers.
đĽ 7 hours ago
Senior DevOps Engineer automating cloud infrastructure, Kubernetes deployments, and CI/CD for Guidehouse government applications. Supporting secure, reliable delivery across development, QA, and operations.
đşđ¸ United States â Remote
đľ $115.2k - $172.8k / year
đ° Grant on 2023-02
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đĽ 10 hours ago
DevSecOps Engineer securing Virta Healthâs cloud-native healthcare platform. Automating application security, IAM, vulnerability management, and compliance across GCP and Kubernetes.
đşđ¸ United States â Remote
đľ $179.5k - $187.9k / year
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đĽ 12 hours ago
Site Reliability Engineer automating secure Azure infrastructure and CI/CD for Arasâs enterprise PLM cloud services. Improving reliability, monitoring, security, and customer environments.
đşđ¸ United States â Remote
đľ $100k - $120k / year
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đĽ 12 hours ago
Senior Site Reliability Engineer building reliable cloud infrastructure for Cross Riverâs fintech products. Driving DevOps, CI/CD, observability, incident response, and operational excellence.
đşđ¸ United States â Remote
đľ $160k - $200k / year
đ° $620M Series D on 2022-03
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor