Senior Site Reliability Engineer

Job not on LinkedIn

🕒 July 28

🏄 California – Remote

infoinfo

💵 $205k - $235k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of The Voleon Group

The Voleon Group

201 - 500 employees

Founded 2007

💸 Finance

🤖 Artificial Intelligence

Finance • Artificial Intelligence

The Voleon Group is an investment management firm that applies machine learning and rigorous statistical research to financial markets. Founded with an academic approach to research, Voleon emphasizes scalable models, risk management, and data-driven financial prediction rather than human intuition. Headquartered near UC Berkeley and operating through Voleon Capital Management LP and affiliates, the company hires Ph. D. -level researchers in statistics, computer science, and related quantitative fields to develop automated investment strategies and manage funds.

📋 Description

• Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies • Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability

🎯 Requirements

• 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead • Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod) • Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.) • Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible) • Experience with cloud infrastructure (AWS or GCP) • Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry) • Experience with distributed storage technologies (Lustre, Ceph, S3) • Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation • Bachelor degree in computer science

🏖️ Benefits

• medical, dental, and vision coverage • life and AD&D insurance • 20 days of paid time off • 9 sick days • 401(k) plan with a company match

Apply Now

Similar Jobs

🕒 July 28

STN Incorporated

11 - 50

🏢 Enterprise

🔒 Cybersecurity

🔧 Hardware

Site Reliability Engineer managing reliability and observability for GPU One platform. Leading incident response and ensuring SLOs for customer SLAs.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 28

Runpod

51 - 200

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🇺🇸 United States – Remote

💵 $150k - $200k / year

💰 $20M Seed on 2024-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 28

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 United States – Remote

💵 $179.4k - $232.1k / year

💰 $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 28

ICF

5001 - 10000

💼 Consulting

🏛️ Government

🏥 Healthcare

DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.

🇺🇸 United States – Remote

💵 $108.5k - $184.4k / year

💰 $29M Grant on 2023-03

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 28

ICF

5001 - 10000

💼 Consulting

🏛️ Government

🏥 Healthcare

Senior DevOps Engineer delivering best in class healthcare reporting services for ICF. Working collaboratively to implement cloud solutions and establish CI/CD pipelines using AWS.

🇺🇸 United States – Remote

💵 $108.5k - $184.4k / year

💰 $29M Grant on 2023-03

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)