Senior Site Reliability Engineer

Vaga não está no LinkedIn

🕒 Julho 28

🏄 California – Remoto

infoinfo

💵 $205.000 - $235.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 10%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of The Voleon Group

The Voleon Group

201 - 500 funcionários

Fundada em 2007

💸 Finanças

🤖 Inteligência Artificial

Finance • Artificial Intelligence

O Voleon Group é uma empresa de gestão de investimentos que aplica aprendizado de máquina e pesquisa estatística rigorosa aos mercados financeiros. Fundada com uma abordagem acadêmica à pesquisa, a Voleon enfatiza modelos escaláveis, gerenciamento de risco e previsões financeiras baseadas em dados, em vez de intuição humana. Com sede próxima à UC Berkeley e operando através da Voleon Capital Management LP e afiliadas, a empresa contrata pesquisadores com doutorado em estatística, ciência da computação e campos quantitativos relacionados para desenvolver estratégias de investimento automatizadas e gerenciar fundos.

Descrição

• Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies • Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability

🎯 Requisitos

• 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead • Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod) • Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.) • Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible) • Experience with cloud infrastructure (AWS or GCP) • Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry) • Experience with distributed storage technologies (Lustre, Ceph, S3) • Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation • Bachelor degree in computer science

🏖️ Benefícios

• medical, dental, and vision coverage • life and AD&D insurance • 20 days of paid time off • 9 sick days • 401(k) plan with a company match

Candidatar-se

Vagas Similares

🕒 Julho 28

STN Incorporated

11 - 50

🏢 Corporativo

🔒 Cibersegurança

🔧 Hardware

Site Reliability Engineer managing reliability and observability for GPU One platform. Leading incident response and ensuring SLOs for customer SLAs.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Runpod

51 - 200

🤖 Inteligência Artificial

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $200.000 / ano

💰 $20.000.000 Seed em 2024-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $179.400 - $232.100 / ano

💰 $75.000.000 Debt Financing - Thumbtack em 2024-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

ICF

5001 - 10000

💼 Consultoria

🏛️ Governo

🏥 Saúde

DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $108.476 - $184.409 / ano

💰 $29.000.000 Grant em 2023-03

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

ICF

5001 - 10000

💼 Consultoria

🏛️ Governo

🏥 Saúde

Senior DevOps Engineer delivering best in class healthcare reporting services for ICF. Working collaboratively to implement cloud solutions and establish CI/CD pipelines using AWS.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $108.476 - $184.409 / ano

💰 $29.000.000 Grant em 2023-03

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório