Search Remote Jobs

Senior AI Tools Engineer, SRE Operations

🔥 0 minutes ago

🏄 California – Remote

infoinfo

💵 $144k - $230k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Build and deploy sophisticated AI-powered tools and products supporting operation and optimization of the global GeForce NOW service • Transform production data streams, including signals, metrics, and logs, into actionable intelligence • Automate root cause analysis for incidents and predict future service trends and patterns • Build and implement AI/ML tools to analyze production data, identify root causes for complex incidents, and identify future operational trends • Lead development of LLM- and Agent-based systems to improve operational efficiency • Establish and maintain data management practices and construct workflows for large-scale data sources vital for model development • Take charge of and enhance LLM-based pipelines while integrating LLM progress into product development • Act as an authority on AI frameworks and recommend platforms, toolsets, and architectural approaches for long-term technical sustainability

🎯 Requirements

• B.S. in Computer Science, Statistics, or Engineering (or equivalent experience) • 5+ years of experience • Strong proficiency in Python • Familiarity with Go or other systems languages is a plus • Practical experience building, optimizing, and deploying AI tools • Strong knowledge of the AI space and current developments, including understanding how LLM-based platforms are built, optimized, and which platforms work best • Hands-on experience with container orchestration (Kubernetes) and cloud environments (AWS cloud) • Active engagement with developments in the AI field and ability to distinguish meaningful advances from noise when making technical decisions • Expertise in automation and handling large-scale data pipelines • Experience applying monitoring and visualization tools, such as Grafana, to interact with data • Excellent ability to handle data sources and pipelines to transform and manage data • Current experience in LLM improvement pipelines and a strong grasp of recent developments in LLM training • Understanding of SRE concepts and experience managing production environments • Experience with Kubernetes, AWS, and other cloud technologies • Excellent knowledge of LLMs and AI Models • Proficiency in automation

🏖️ Benefits

• Competitive salary package • Equity • Benefits

Apply Now

Similar Jobs

🔥 5 hours ago

Koniag Government Services

1001 - 5000

🏛️ Government

🎖️ Defense

💼 Consulting

Senior DevOps Engineer securing AWS/Azure cloud infrastructure for Koniag Government Services. Automating DevSecOps, CI/CD security, compliance, and incident response for federal customers.

🔥 7 hours ago

Guidehouse

10,000+ employees

🏥 Healthcare

🎖️ Defense

📦 Logistics

Senior DevOps Engineer automating cloud infrastructure, Kubernetes deployments, and CI/CD for Guidehouse government applications. Supporting secure, reliable delivery across development, QA, and operations.

🔥 10 hours ago

Virta Health

201 - 500

🏥 Healthcare

⚕️ Healthcare Insurance

🧘 Wellness

DevSecOps Engineer securing Virta Health’s cloud-native healthcare platform. Automating application security, IAM, vulnerability management, and compliance across GCP and Kubernetes.

🔥 12 hours ago

Aras Corporation

501 - 1000

🏭 Manufacturing

💼 Consulting

📦 Logistics

Site Reliability Engineer automating secure Azure infrastructure and CI/CD for Aras’s enterprise PLM cloud services. Improving reliability, monitoring, security, and customer environments.

🔥 12 hours ago

Cross River

501 - 1000

🏦 Banking

💳 Fintech

☁️ SaaS

Senior Site Reliability Engineer building reliable cloud infrastructure for Cross River’s fintech products. Driving DevOps, CI/CD, observability, incident response, and operational excellence.