Search Remote Jobs

Senior Site Reliability Engineer

Job not on LinkedIn

πŸ”₯ 9 minutes ago

Apply Now
Find Similar Remote Jobs

πŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of The Voleon Group

The Voleon Group

201 - 500 employees

Founded 2007

πŸ’Έ Finance

πŸ€– Artificial Intelligence

Finance β€’ Artificial Intelligence

The Voleon Group is an investment management firm that applies machine learning and rigorous statistical research to financial markets. Founded with an academic approach to research, Voleon emphasizes scalable models, risk management, and data-driven financial prediction rather than human intuition. Headquartered near UC Berkeley and operating through Voleon Capital Management LP and affiliates, the company hires Ph. D. -level researchers in statistics, computer science, and related quantitative fields to develop automated investment strategies and manage funds.

πŸ“‹ Description

β€’ Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise β€’ Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability β€’ Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams β€’ Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do β€’ Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies β€’ Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability

🎯 Requirements

β€’ 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead β€’ Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod) β€’ Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.) β€’ Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible) β€’ Experience with cloud infrastructure (AWS or GCP) β€’ Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry) β€’ Experience with distributed storage technologies (Lustre, Ceph, S3) β€’ Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation β€’ Bachelor degree in computer science

πŸ–οΈ Benefits

β€’ medical, dental, and vision coverage β€’ life and AD&D insurance β€’ 20 days of paid time off β€’ 9 sick days β€’ 401(k) plan with a company match

Apply Now

Similar Jobs

πŸ”₯ 20 minutes ago

STN Incorporated

11 - 50

🏒 Enterprise

πŸ”’ Cybersecurity

πŸ”§ Hardware

Site Reliability Engineer managing reliability and observability for GPU One platform. Leading incident response and ensuring SLOs for customer SLAs.

πŸ”₯ 29 minutes ago

Zafran Security

51 - 200

πŸ” Security

Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.

πŸ”₯ 32 minutes ago

Runpod

51 - 200

πŸ€– Artificial Intelligence

☁️ SaaS

🀝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

πŸ”₯ 53 minutes ago

Talkiatry

501 - 1000

πŸ₯ Healthcare

πŸ‘₯ B2C

Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.

πŸ”₯ 57 minutes ago

Careerswift

2 - 10

πŸ‘₯ HR Tech

🎯 Recruiter

☁️ SaaS

DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.