Principal Operations Engineer, Reliability

🔥 42 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of FluidStack

FluidStack

11 - 50 employees

🤖 Artificial Intelligence

Artificial Intelligence • Cloud Computing

FluidStack is a company that provides GPU supercomputing infrastructure for AI labs. It offers on-demand access to thousands of Nvidia GPUs, enabling large-scale AI training and inference. The company specializes in deploying and managing large GPU clusters with support for technologies like Kubernetes and Slurm, ensuring high availability and excellent support. FluidStack provides a fully managed cloud infrastructure, helping AI companies to focus on developing models without worrying about the underlying hardware. They emphasize performance and cost-efficiency, offering services that scale to thousands of GPUs with high uptime and rapid response times.

📋 Description

• Own fleet reliability engineering: define availability targets, measure them honestly, and close the gap. • Run root cause analysis on the fleet's worst incidents and drive corrective actions to done across every site. • Build the failure data pipeline, facility and hardware both, that turns incident history into engineering priorities. • Set the maintenance strategy (reliability-centered, condition-based) so the fleet spends effort where the failure data says to.

🎯 Requirements

• You've owned reliability for critical infrastructure and moved the availability number, not just reported it. • You've led root cause analyses that found the real cause, not the convenient one. • You work fluently with failure data: Weibull, Pareto, and FMEA are tools you actually use, not terms you know. • You get corrective actions closed across teams you don't manage. • Bonus: Data center or power generation reliability. Liquid cooling systems. CMMS analytics. CRE or CMRP certification.

🏖️ Benefits

• base salary • equity for all full time roles • benefits • commissions plans if applicable

Apply Now

Similar Jobs

🔥 2 hours ago

Tavily

51 - 200

🔌 API

🤖 Artificial Intelligence

☁️ SaaS

Head of Forward Deployed Engineering at Tavily, managing strategic customer engagements and leading a growing engineering team. Overseeing technical strategy for enterprise accounts across product and engineering teams.

🇺🇸 United States – Remote

💵 $230k - $310k / year

🔥 Funding within the last year

💰 $20M Series A - Tavily on 2025-08

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

Distributed Systems

Python

🔥 4 hours ago

Dayforce

5001 - 10000

👥 HR Tech

☁️ SaaS

🏢 Enterprise

Maintenance & Reliability Engineer leading enterprise-wide initiatives to enhance asset performance and reliability in a manufacturing environment. Collaborating with engineering and maintenance teams to improve operational efficiency.

🇺🇸 United States – Remote

💵 $115k - $150k / year

💰 $1G Post-IPO Debt - Dayforce on 2024-03

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 4 hours ago

Cadence Solutions

11 - 50

💼 Consulting

🏢 Enterprise

🤝 B2B

Staff DevOps Engineer responsible for the cloud infrastructure powering care delivery for over 100,000 patients. Cadence Solutions automates chronic disease care with clinical AI technology.

AWS

Cloud

Kubernetes

Terraform

🔥 6 hours ago

Latitude IT Solutions | SDVOSB

11 - 50

🤝 B2B

💼 Consulting

🏛️ Government

DevSecOps & SRE Technical Director leading CI/CD pipelines and security automation for a large-scale VA health IT platform. Managing deployments, infrastructure provisioning, and compliance controls.

AWS

Cloud

Kafka

Kubernetes

Terraform

Vault

🔥 7 hours ago

Counterpart Health

51 - 200

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

Director of Site Reliability Engineering at Counterpart Health managing a team of SREs and enhancing infrastructure reliability. Leading strategic initiatives across multiple regions and supporting healthcare innovation.

Cloud

Google Cloud Platform

Grafana

Kubernetes

Postgres

Prometheus

Python

SQL

Terraform

Go