Search Remote Jobs

Senior AI Tools Engineer, SRE Operations

🔥 2 hours ago

🏄 California – Remote

info

đź’µ $144k - $230k / year

⏰ Full Time

đźź  Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

đź“‹ Description

• Build and deploy sophisticated AI-powered tools and products supporting operation and optimization of the global GeForce NOW service • Transform production data streams, including signals, metrics, and logs, into actionable intelligence • Automate root cause analysis for incidents and predict future service trends and patterns • Build and implement AI/ML tools to identify incident root causes and operational trends • Lead development of LLM- and agent-based systems to improve operational efficiency • Establish and maintain data management practices and pipelines for large-scale model-development data • Own and enhance LLM-based pipelines • Integrate knowledge of LLM progress into product development • Serve as an authority on AI frameworks and recommend platforms, toolsets, and architectures for long-term product sustainability

🎯 Requirements

• B.S. in Computer Science, Statistics, or Engineering, or equivalent experience • 5+ years of experience • Strong proficiency in Python • Familiarity with Go or other systems languages is a plus • Practical experience building, optimizing, and deploying AI tools • Strong knowledge of AI developments and LLM-based platforms • Hands-on experience with Kubernetes and AWS cloud environments • Ability to distinguish meaningful AI advances from noise in technical decisions • Expertise in automation and large-scale data pipelines • Experience with monitoring and visualization tools such as Grafana • Excellent ability to handle, transform, and manage data sources and pipelines • Understanding of SRE principles and production-environment management is preferred • Experience with LLM improvement pipelines and recent LLM training developments is preferred • Knowledge of LLMs and AI models, including ability to recommend sustainable platforms and approaches

🏖️ Benefits

• Equity • Benefits • Competitive salary package • Inclusive and supportive work environment

Apply Now

Similar Jobs

🔥 3 hours ago

SitusAMC

5001 - 10000

đź’Ľ Consulting

📦 Logistics

🏠 Real Estate

Site Reliability Engineer operating AWS cloud infrastructure for SitusAMC’s real estate technology solutions. Improving reliability, automation, observability, security, and application migrations.

🔥 5 hours ago

GE Vernova

10,000+ employees

đź’Ľ Consulting

📦 Logistics

🏭 Manufacturing

Senior Reliability Engineer improving embedded protection, control, and software products for GE Vernova’s decarbonization mission. Driving testing, field analytics, grid reliability, and cybersecurity compliance.

🔥 5 hours ago

SailPoint

1001 - 5000

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Senior Staff DevOps Engineer scaling AWS Kubernetes infrastructure for SailPoint’s identity security platform. Leading enterprise service mesh adoption, PCI-compliant operations, and cloud-native reliability across global teams.

🔥 6 hours ago

Pacvue

501 - 1000

đź’Ľ Consulting

📣 Marketing

📦 Logistics

Senior DevOps Engineer building AWS/Kubernetes infrastructure and CI/CD systems for Pacvue’s commerce media platform. Improving reliability, security, observability, and developer productivity across engineering teams.

🔥 8 hours ago

Affirm

1001 - 5000

đź’ł Fintech

👥 B2C

🛍️ eCommerce

Senior backend engineer building Affirm’s reliability platform for scalable buy-now-pay-later systems. Designing distributed services, AI-assisted tooling, and operational safeguards for production health.