
11 - 50 employees
🤖 Artificial Intelligence
Artificial Intelligence • Cloud Computing
FluidStack is a company that provides GPU supercomputing infrastructure for AI labs. It offers on-demand access to thousands of Nvidia GPUs, enabling large-scale AI training and inference. The company specializes in deploying and managing large GPU clusters with support for technologies like Kubernetes and Slurm, ensuring high availability and excellent support. FluidStack provides a fully managed cloud infrastructure, helping AI companies to focus on developing models without worrying about the underlying hardware. They emphasize performance and cost-efficiency, offering services that scale to thousands of GPUs with high uptime and rapid response times.
🔥 17 hours ago
🇺🇸 United States – Remote
💵 $218k - $263k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
🤖 Artificial Intelligence
Artificial Intelligence • Cloud Computing
FluidStack is a company that provides GPU supercomputing infrastructure for AI labs. It offers on-demand access to thousands of Nvidia GPUs, enabling large-scale AI training and inference. The company specializes in deploying and managing large GPU clusters with support for technologies like Kubernetes and Slurm, ensuring high availability and excellent support. FluidStack provides a fully managed cloud infrastructure, helping AI companies to focus on developing models without worrying about the underlying hardware. They emphasize performance and cost-efficiency, offering services that scale to thousands of GPUs with high uptime and rapid response times.
• Own the full lifecycle of physical installation projects: coordinate LVC schedules, run QA/QC on contractor work, and ensure installs meet design standards, safety codes, and labeling policy. • Serve as the primary on-site technical point of contact for the Low Voltage Contractor and internal engineering teams, and act as remote hands during deployment and validation. • Manage the LVC deployment schedule to milestones, coordinate access, material delivery, and staging with facility operations, and report progress and delays. • Run multi-stage QA inspections on rack placement, racking, cable routing, grounding, and labeling against the BoM and design documentation. • Conduct or oversee optical loss testing (OLTS), OTDR analysis, and insertion/return loss on all fiber links, and verify copper (Cat6a/Cat8) integrity with certified testers. • Act as the SME for fiber and structured cabling best practices (BICSI), generate final test reports, as-builts, and RCAs, and own closeout and handoff to operations.
• You're an SME in fiber optics, structured cabling, and data center physical infrastructure, with hands-on deployment on mission-critical sites. • You run advanced fiber testing (OTDR, VFL, power meter) and know installation techniques and design standards cold. • You've racked, stacked, and cabled servers, network, and storage hardware at scale. • You read and translate technical drawings and schematics (AutoCAD, Visio), and drive projects to completion through ambiguity. • You'll travel 50% for onsite turn-ups, wherever the work is. • Bonus: BICSI RCDD or vendor certifications (Cisco, Juniper). Hyperscale standardization programs. ITIL.
• Competitive total compensation package (salary + equity). • Retirement or pension plan, in line with local norms. • Health, dental, and vision insurance. • Generous PTO policy, in line with local norms.
Apply Now🕒 3 days ago
DevOps Engineer IV for Envision's Engineering team. Collaborating on implementing advanced CI/CD pipelines and infrastructure as code.
🇺🇸 United States – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Airflow
Amazon Redshift
AWS
Cloud
Docker
ETL
Jenkins
Kubernetes
Python
Spark
SQL
SSIS
Terraform
🕒 3 days ago
Corporate Reliability Engineer enhancing reliability and performance at Arclin's manufacturing processes. Involves data analysis, reliability assessments, and cross-functional collaboration for continuous improvement.
🕒 3 days ago
Senior Lead Database Reliability Engineer ensuring database reliability and operational excellence for a sports betting platform. Leading technical roadmap and optimizing performance across various systems.
🇺🇸 United States – Remote
💵 $168k - $210k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Cloud
Kubernetes
MongoDB
MySQL
Postgres
Python
Redis
SQL
Terraform
Go
🕒 3 days ago
Senior Site Reliability Engineer responsible for infrastructure managing high-volume logistics operations on GCP. Collaborating to enhance reliability, automate processes, and improve monitoring practices.
🇺🇸 United States – Remote
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Cloud
Distributed Systems
Docker
Google Cloud Platform
Kubernetes
Python
Terraform
TypeScript
Go
🕒 3 days ago
Software Engineer Manager leading Reliability Engineering practices for Home Depot's cloud foundation. Ensuring resilience, performance, and security through extensive automation and incident management.
🇺🇸 United States – Remote
💵 $140k - $240k / year
💰 Debt Financing on 2007-07
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Cloud
Java
Kubernetes
Terraform