Principal Operations Engineer, Reliability

🕒 5 days ago

🇺🇸 United States – Remote

💵 $220k - $260k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of FluidStack

FluidStack

11 - 50 employees

🤖 Artificial Intelligence

Artificial Intelligence • Cloud Computing

FluidStack is a company that provides GPU supercomputing infrastructure for AI labs. It offers on-demand access to thousands of Nvidia GPUs, enabling large-scale AI training and inference. The company specializes in deploying and managing large GPU clusters with support for technologies like Kubernetes and Slurm, ensuring high availability and excellent support. FluidStack provides a fully managed cloud infrastructure, helping AI companies to focus on developing models without worrying about the underlying hardware. They emphasize performance and cost-efficiency, offering services that scale to thousands of GPUs with high uptime and rapid response times.

📋 Description

• Own fleet reliability engineering: define availability targets, measure them honestly, and close the gap. • Run root cause analysis on the fleet's worst incidents and drive corrective actions to done across every site. • Build the failure data pipeline, facility and hardware both, that turns incident history into engineering priorities. • Set the maintenance strategy (reliability-centered, condition-based) so the fleet spends effort where the failure data says to.

🎯 Requirements

• You've owned reliability for critical infrastructure and moved the availability number, not just reported it. • You've led root cause analyses that found the real cause, not the convenient one. • You work fluently with failure data: Weibull, Pareto, and FMEA are tools you actually use, not terms you know. • You get corrective actions closed across teams you don't manage. • Bonus: Data center or power generation reliability. Liquid cooling systems. CMMS analytics. CRE or CMRP certification.

🏖️ Benefits

• base salary • equity for all full time roles • benefits • commissions plans if applicable

Apply Now

Similar Jobs

🕒 5 days ago

Tavily

51 - 200

🔌 API

🤖 Artificial Intelligence

☁️ SaaS

Head of Forward Deployed Engineering at Tavily, managing strategic customer engagements and leading a growing engineering team. Overseeing technical strategy for enterprise accounts across product and engineering teams.

🇺🇸 United States – Remote

💵 $230k - $310k / year

🔥 Funding within the last year

💰 $20M Series A - Tavily on 2025-08

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 5 days ago

Cadence Solutions

11 - 50

💼 Consulting

🏢 Enterprise

🤝 B2B

Staff DevOps Engineer responsible for the cloud infrastructure powering care delivery for over 100,000 patients. Cadence Solutions automates chronic disease care with clinical AI technology.

🇺🇸 United States – Remote

💵 $200k - $260k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 5 days ago

Counterpart Health

51 - 200

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

Director of Site Reliability Engineering at Counterpart Health managing a team of SREs and enhancing infrastructure reliability. Leading strategic initiatives across multiple regions and supporting healthcare innovation.

🇺🇸 United States – Remote

💵 $187k - $243k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 5 days ago

Finalsite

201 - 500

📚 Education

☁️ SaaS

🤝 B2B

Staff Site Reliability Engineer leading Finalsite's infrastructure evolution and operational excellence practices. Collaborating with engineering leadership to enhance CI/CD and multi-cloud reliability.

🇺🇸 United States – Remote

💵 $180k - $250k / year

💰 Debt financing on 2014-12

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 5 days ago

Whitespace

11 - 50

🎖️ Defense

🏛️ Government

🤖 Artificial Intelligence

Senior DevSecOps Engineer enhancing cybersecurity compliance for federal standards and DoD authorization processes. Leading secure CI/CD implementations and DevSecOps toolchain management for government projects.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)