Senior Site Reliability Engineer – SRE

Vaga não está no LinkedIn

🕒 Março 18

🌽 Illinois – Remoto

infoinfo

💵 $165.000 - $225.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 54%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Moonlite

Moonlite

1 - 10 funcionários

📚 Educação

🏪 Marketplace

👥 B2C

Education • Marketplace • B2C

Moonlite é uma plataforma web orientada pela comunidade que ajuda as pessoas a descobrir, comparar e confiar em formas comprovadas de ganhar dinheiro. Ela organiza milhares de ideias de negócios, recursos, criadores e cursos, todos validados e avaliados por usuários reais, para que os interessados possam evitar exageros e focar no que realmente funciona. A Moonlite oferece discussões em comunidade, avaliações, comparações lado a lado e uma pesquisa rápida para conectar os usuários com caminhos de renda adequados, com o objetivo de ajudar os indivíduos a construir liberdade financeira com confiança.

Descrição

• Design, build, and operate production Kubernetes clusters on bare-metal infrastructure. • Implement and operate custom Kubernetes networking solutions. • Develop and maintain custom Kubernetes operators and controllers. • Deploy and optimize NVIDIA GPU operators and custom scheduling logic for GPU workloads. • Build deep integrations between Kubernetes and underlying infrastructure. • Design and implement automation using Terraform, Ansible, Helm, and custom operators. • Manage production bare-metal infrastructure across multiple regions ensuring high availability, fault tolerance, and graceful degradation. • Build comprehensive monitoring, logging, and alerting using Prometheus, Grafana, and ELK stack. • Identify and resolve performance bottlenecks across infrastructure domains.

🎯 Requisitos

• 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at scale. • Deep hands-on experience building and operating production Kubernetes clusters on bare-metal infrastructure. • Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. • Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments. • Proficiency with infrastructure-as-code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead. • Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production. • Experience building and maintaining comprehensive monitoring solutions using tools like Prometheus, Grafana, and centralized logging systems. • Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems. • Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency. • Demonstrated ability to troubleshoot complex issues under pressure, manage incidents effectively, and communicate clearly during outages. • Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.

🏖️ Benefícios

• 6% 401(k) match • Fully covered health insurance premiums • Other comprehensive offerings to support your well-being and success as we grow together.

Candidatar-se

Vagas Similares

🕒 Março 18

Vytwo Technologies Inc

201 - 500

💼 Consultoria

📦 Logística

🎯 Recrutamento

Meanstack Architect with DevOps expertise for TCoE, designing scalable applications and leading technical teams in a fully remote environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $45 - $50 / hora

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 13

Resolve Tech Solutions

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

DevOps Lead Engineer at RTS responsible for scalable cloud infrastructure design and CI/CD pipeline optimization. Collaborating across teams to drive automation, governance, and cost optimization.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 12

Keeper Security, Inc.

501 - 1000

🔒 Cibersegurança

☁️ SaaS

🏢 Corporativo

Senior DevOps Engineer managing IL5-compliant infrastructure for Keeper Security, working in high-security environments and collaborating with various engineering teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Private Equity Round - Keeper Security em 2023-05

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 10

Deepgram

51 - 200

💼 Consultoria

🏥 Saúde

📦 Logística

Site Reliability Engineer managing AI/ML infrastructure for Deepgram. Architecting, building, and optimizing hybrid systems with Kubernetes, AWS, and Terraform.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $220.000 / ano

💰 $47.000.000 Series B em 2022-11

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 7

Inetum

10.000+ funcionários

💼 Consultoria

🏥 Saúde

🛡️ Seguros

Expert DevOps / DevSecOps supporting Generative AI initiatives at Inetum for digital transformation in the United States. Designing high-value GenAI use cases and integrating new tools and practices.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Post-IPO Equity em 2007-03

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇫🇷 Francês obrigatório