Senior Site Reliability Engineer – SRE

Vaga não está no LinkedIn

🕒 Março 18

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Moonlite

Moonlite

1 - 10 funcionários

📚 Educação

🏪 Marketplace

👥 B2C

Education • Marketplace • B2C

Moonlite é uma plataforma web orientada pela comunidade que ajuda as pessoas a descobrir, comparar e confiar em formas comprovadas de ganhar dinheiro. Ela organiza milhares de ideias de negócios, recursos, criadores e cursos, todos validados e avaliados por usuários reais, para que os interessados possam evitar exageros e focar no que realmente funciona. A Moonlite oferece discussões em comunidade, avaliações, comparações lado a lado e uma pesquisa rápida para conectar os usuários com caminhos de renda adequados, com o objetivo de ajudar os indivíduos a construir liberdade financeira com confiança.

Descrição

• Design, build, and operate production Kubernetes clusters on bare-metal infrastructure. • Implement and operate custom Kubernetes networking solutions. • Develop and maintain custom Kubernetes operators and controllers. • Deploy and optimize NVIDIA GPU operators and custom scheduling logic for GPU workloads. • Build deep integrations between Kubernetes and underlying infrastructure. • Design and implement automation using Terraform, Ansible, Helm, and custom operators. • Manage production bare-metal infrastructure across multiple regions ensuring high availability, fault tolerance, and graceful degradation. • Build comprehensive monitoring, logging, and alerting using Prometheus, Grafana, and ELK stack. • Identify and resolve performance bottlenecks across infrastructure domains.

🎯 Requisitos

• 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at scale. • Deep hands-on experience building and operating production Kubernetes clusters on bare-metal infrastructure. • Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. • Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments. • Proficiency with infrastructure-as-code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead. • Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production. • Experience building and maintaining comprehensive monitoring solutions using tools like Prometheus, Grafana, and centralized logging systems. • Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems. • Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency. • Demonstrated ability to troubleshoot complex issues under pressure, manage incidents effectively, and communicate clearly during outages. • Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.

🏖️ Benefícios

• 6% 401(k) match • Fully covered health insurance premiums • Other comprehensive offerings to support your well-being and success as we grow together.

Candidatar-se

Vagas Similares

🕒 Março 18

Vytwo Technologies Inc

201 - 500

💼 Consultoria

📦 Logística

🎯 Recrutamento

Meanstack Architect with DevOps expertise for TCoE, designing scalable applications and leading technical teams in a fully remote environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $45 - $50 / hora

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 17

Truv

51 - 200

💼 Consultoria

🏥 Saúde

📦 Logística

Senior DevOps Engineer architecting and scaling AWS infrastructure and building observability platforms. Leading compliance projects and optimizing CI/CD pipelines in a remote setup.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 14

Panopto

51 - 200

☁️ SaaS

📚 Educação

🏢 Corporativo

Mid-Level DevOps Engineer at Panopto transforming outdated build processes into automated pipelines. Elevate the engineering experience by enhancing delivery lifecycle and collaboration.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $155.000 - $175.000 / ano

💰 Private Equity Round em 2021-04

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 13

Sword Health

201 - 500

🏥 Saúde

💼 Consultoria

📦 Logística

DevOps Engineer at Sword Health designing scalable infrastructure and automating processes to enhance AI healthcare solutions. Collaborating with cross-functional teams in a remote work environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $140.000 - $220.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 13

Resolve Tech Solutions

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

DevOps Lead Engineer at RTS responsible for scalable cloud infrastructure design and CI/CD pipeline optimization. Collaborating across teams to drive automation, governance, and cost optimization.

🗣️🇺🇸🇬🇧 Inglês obrigatório