Senior Site Reliability Engineer

Stelle nicht auf LinkedIn

🕒 vor 1 Monat

🏄 California – Remote

infoinfo

💵 $205.000 - $235.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 4%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of The Voleon Group

The Voleon Group

201 - 500 Mitarbeiter

Gegründet 2007

💸 Finanzen

🤖 Künstliche Intelligenz

Finance • Artificial Intelligence

Die Voleon Gruppe ist ein Investment-Management-Unternehmen, das maschinelles Lernen und rigorose statistische Forschung auf Finanzmärkte anwendet. Gegründet mit einem akademischen Ansatz zur Forschung, legt Voleon Wert auf skalierbare Modelle, Risikomanagement und datengetriebene finanzielle Vorhersagen anstelle menschlicher Intuition. Der Hauptsitz befindet sich in der Nähe der UC Berkeley und das Unternehmen operiert über Voleon Capital Management LP und angeschlossene Unternehmen. Es werden Forscher mit Doktortitel in Statistik, Informatik und verwandten quantitativen Bereichen eingestellt, um automatisierte Investitionsstrategien zu entwickeln und Fonds zu verwalten.

Beschreibung

• Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies • Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability

🎯 Anforderungen

• 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead • Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod) • Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.) • Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible) • Experience with cloud infrastructure (AWS or GCP) • Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry) • Experience with distributed storage technologies (Lustre, Ceph, S3) • Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation • Bachelor degree in computer science

🏖️ Vorteile

• medical, dental, and vision coverage • life and AD&D insurance • 20 days of paid time off • 9 sick days • 401(k) plan with a company match

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

STN Incorporated

11 - 50

🏢 Unternehmen

🔒 Cybersecurity

🔧 Hardware

Site Reliability Engineer managing reliability and observability for GPU One platform. Leading incident response and ensuring SLOs for customer SLAs.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Zafran Security

51 - 200

🔐 Sicherheit

Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Runpod

51 - 200

🤖 Künstliche Intelligenz

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🇺🇸 Vereinigte Staaten – Remote

💵 $150.000 - $200.000 / Jahr

💰 €20.000.000 Seed im 2024-06

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Multi Media, LLC

51 - 200

💼 Beratung

📣 Marketing

📱 Medien

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

🇺🇸 Vereinigte Staaten – Remote

💵 $169.000 - $215.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Thumbtack

1001 - 5000

🏪 Marktplatz

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 Vereinigte Staaten – Remote

💵 $179.400 - $232.100 / Jahr

💰 €75.000.000 Debt Financing - Thumbtack im 2024-07

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich