Senior/Staff Site Reliability Engineer – Data Center

🔥 0 minutes ago

🍂 Massachusetts – Remote

info

💵 $165.8k - $224.4k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of PathAI

PathAI

501 - 1000 employees

Founded 2016

🏥 Healthcare

💼 Consulting

📦 Logistics

💰 $165M Series C on 2021-05

Healthcare • Consulting • Logistics

PathAI is a leading healthcare technology company dedicated to advancing pathology through the use of artificial intelligence. Their mission is to improve patient outcomes by providing AI-powered technology that offers valuable insights for biomarker discovery and drug development. PathAI collaborates closely with biopharma and pathology laboratories to enhance laboratory workflows and improve diagnostics. Their AI-driven pathology solutions, such as the AISight Digital Pathology Image Management System, are utilized by major laboratories and research centers worldwide, helping to power digital pathology and precision medicine initiatives. PathAI's technology is leveraged by top biopharma companies to transform drug discovery and diagnostics, making significant contributions to the fields of healthcare and biotechnology.

📋 Description

• Advance operations by implementing SRE best practices focused on users, monitoring, and automation • Design, build, and operate the data center supporting the Machine Learning team • Build highly secure on-premises environments handling NIST/ISO standards • Integrate on-premises data center environments with existing cloud infrastructure to create a seamless hybrid cloud environment • Improve infrastructure reliability and resilience through root-cause analysis, design-gap reviews, and implementation improvements • Participate in platform on-call rotations and assist with urgent incident response

🎯 Requirements

• 8+ years of relevant experience • Familiarity with modern data center network designs and operating across network layers • Experience administering physical hardware stacks in production settings, including iDRAC, IPMI, Nvidia UFM, and Juniper systems • Experience with virtualization, containerization, or container orchestration platforms, including EKS-Anywhere, ClusterAPI, or KVM • Knowledge of storage solutions and optimization for high-performance workloads, including Quobyte, S3, FSx, or EFS • Automation experience using scripting and configuration management tools such as Ansible and RedFish • Experience building monitoring infrastructure with Datadog, Grafana, or Prometheus • Experience managing critical production infrastructure, incident response, scaling, and rapid-growth challenges • Bachelor's degree in Computer Science or equivalent experience • Ability to learn quickly in a complex environment • Occasional travel to onsite data center locations

🏖️ Benefits

• Annual pay range of $165,750 - $224,450 • Remote work option • On-target commission for eligible roles (the posting states the cash compensation framework includes it, though this role's commission eligibility is not explicitly confirmed)

Apply Now

Similar Jobs

🔥 13 minutes ago

Ardent

51 - 200

💼 Consulting

🎖️ Defense

📦 Logistics

DevSecOps Engineer securing cloud products and services for Ardent’s federal national security and defense missions. Automating deployments, vulnerability mitigation, and enterprise system architecture.

🔥 51 minutes ago

Hexion Inc.

1001 - 5000

🚘 Automotive

🏗️ Construction

🏭 Manufacturing

Reliability Engineer improving asset performance across Hexion’s North American manufacturing plants. Leading failure elimination, maintenance optimization, and cross-site reliability standardization.

🔥 3 hours ago

Vontier

5001 - 10000

🚘 Automotive

⚡ Energy

🔧 Hardware

Hardware Reliability Engineer leading reliability planning, validation, and failure analysis for Gilbarco Veeder-Root fueling equipment. Improving durability, serviceability, and field performance across complex electromechanical products.

🔥 3 hours ago

Hexion Inc.

1001 - 5000

🚘 Automotive

🏗️ Construction

🏭 Manufacturing

Reliability Engineer reducing downtime and improving asset performance across Hexion’s North American manufacturing plants. Leading RCA, maintenance strategy, CMMS execution, KPI analysis, and cross-site reliability standardization.

🔥 9 hours ago

Salve.Inno

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior SRE building and operating highly available cloud platforms for Salve.Inno Consulting clients. Improving observability, automation, incident response, and production reliability across engineering teams.