
201 - 500 employees
Founded 2007
πΈ Finance
π€ Artificial Intelligence
Finance β’ Artificial Intelligence
The Voleon Group is an investment management firm that applies machine learning and rigorous statistical research to financial markets. Founded with an academic approach to research, Voleon emphasizes scalable models, risk management, and data-driven financial prediction rather than human intuition. Headquartered near UC Berkeley and operating through Voleon Capital Management LP and affiliates, the company hires Ph. D. -level researchers in statistics, computer science, and related quantitative fields to develop automated investment strategies and manage funds.
π₯ 9 minutes ago
π California β Remote
π΅ $205k - $235k / year
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
Founded 2007
πΈ Finance
π€ Artificial Intelligence
Finance β’ Artificial Intelligence
The Voleon Group is an investment management firm that applies machine learning and rigorous statistical research to financial markets. Founded with an academic approach to research, Voleon emphasizes scalable models, risk management, and data-driven financial prediction rather than human intuition. Headquartered near UC Berkeley and operating through Voleon Capital Management LP and affiliates, the company hires Ph. D. -level researchers in statistics, computer science, and related quantitative fields to develop automated investment strategies and manage funds.
β’ Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise β’ Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability β’ Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams β’ Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do β’ Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies β’ Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability
β’ 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead β’ Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod) β’ Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.) β’ Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible) β’ Experience with cloud infrastructure (AWS or GCP) β’ Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry) β’ Experience with distributed storage technologies (Lustre, Ceph, S3) β’ Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation β’ Bachelor degree in computer science
β’ medical, dental, and vision coverage β’ life and AD&D insurance β’ 20 days of paid time off β’ 9 sick days β’ 401(k) plan with a company match
Apply Nowπ₯ 20 minutes ago
Site Reliability Engineer managing reliability and observability for GPU One platform. Leading incident response and ensuring SLOs for customer SLAs.
π₯ 29 minutes ago
Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.
π₯ 32 minutes ago
Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.
πΊπΈ United States β Remote
π΅ $150k - $200k / year
π° $20M Seed on 2024-06
β° Full Time
π‘ Mid-level
π Senior
β DevOps & Site Reliability Engineer (SRE)
π₯ 53 minutes ago
Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.
πΊπΈ United States β Remote
π΅ $160k - $185k / year
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)
π₯ 57 minutes ago
DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.