Site Reliability Engineer – Engineering Productivity

🕒 Setembro 2

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 10%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Arista Networks

Arista Networks

1001 - 5000 funcionários

Fundada em 2004

🏢 Corporativo

📡 Telecomunicações

💰 $2.600.000 Post-IPO Debt em 2015-05

Enterprise • Telecommunications • Cloud

Arista Networks é líder na construção de redes escaláveis, de alto desempenho e com latência ultrabaixa para ambientes modernos de data centers e computação em nuvem. A empresa oferece uma ampla gama de soluções de networking, incluindo o Arista Extensible Operating System (EOS) para cloud networking, o CloudVision para automação e visibilidade de redes, e uma variedade de switches e plataformas de roteamento de alto desempenho. As soluções da Arista são projetadas para redes WAN empresariais, roteamento em nível de nuvem, segmentação multi-domínio, segurança e modelos modernos de operação de rede, tornando-se uma escolha crucial para data centers de grande escala e ambientes de nuvem.

Descrição

• Build, safely and incrementally deploy, and operate critical production systems • Focus on scalability, reliability, observability, performance, and security • Monitor, support, and enhance developer experience across services • Build automation to remove toil and operate production systems efficiently • Monitor and respond to alerts, enhance alerting, and establish automated alert handling • Create and maintain incident-response runbooks • Build and deploy new systems with scalability, reliability, and observability as primary requirements • Triage platform and infrastructure issues and assist Arista software engineers with triage • Engage with third-party vendor support • Deploy new systems in a staged manner • Write postmortems and develop solutions to prevent recurring incidents • Plan and communicate production-system maintenance windows • Identify infrastructure issues causing workflow bottlenecks and limitations for product-development teams • Design and implement solutions to resolve infrastructure issues • Survey and adopt infrastructure and platform best practices • Implement fault tolerance and performance improvements to increase system availability • Study open-source system designs and implementation details to improve triage and fix resolution

🎯 Requisitos

• At least BSc Computer Science or Engineering + 5 years’ experience, MS Computer Science or Engineering + 5 years’ experience, or equivalent work experience • Knowledge of one or more of Go, Python, or shell scripting to implement medium-complexity automation workflows • Knowledge of Linux or UNIX from an administration and debugging perspective • Hands-on experience operating software systems at scale • Experience in server provisioning, especially from storage and networking perspectives • Strong problem-solving and software troubleshooting skills • Experience with infrastructure-as-code • Experience managing databases such as MariaDB, PostgreSQL, or MongoDB • Experience with Docker and virtualization technologies such as KVM, QEMU, or Kata Containers • Experience managing monitoring stacks such as Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos • Experience managing Elasticsearch clusters • Experience managing Artifactory or Docker registries • Experience managing CI/CD systems such as ArgoCD or Spinnaker • Experience managing version-control systems such as Perforce or Gerrit • Experience with infrastructure-as-code frameworks such as Ansible • Experience managing large Java applications • Experience in storage infrastructure management such as NAS, SAN, or Ceph

🏖️ Benefícios

• Great Place to Work awards for engineering, diversity, compensation, and work-life balance • Engineers have complete ownership of their projects • Flat and streamlined management structure • Opportunities to work across various domains • Access to every part of the company • Inclusive environment valuing diversity of thought and perspectives

Candidatar-se

Vagas Similares

🕒 Setembro 1

Miratech

501 - 1000

🤝 B2B

💼 Consultoria

☁️ SaaS

Platform DevOps Engineer managing AWS infrastructure, Kubernetes, and CI/CD pipelines for Miratech’s global IT services. Automating reliable cloud operations with Python, Terraform, Ansible, and DevOps tooling.

🇮🇳 Índia – Remoto

💰 Private Equity Round em 2022-04

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 1

Playpower Labs

11 - 50

📚 Educação

🤖 Inteligência Artificial

DevOps Engineer running AWS infrastructure, CI/CD, security, and reliability for PlayPower Labs’ EdTech software. Supporting products used by millions of students and teachers.

🇮🇳 Índia – Remoto

💰 $120.000 Pre Seed Round em 2013-05

⏰ Tempo Integral

🟢 Júnior

🟡 Pleno

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🚫👨‍🎓 Sem graduação necessária

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 31

Hyland

1001 - 5000

🤝 B2B

☁️ SaaS

🏢 Corporativo

Senior DevOps Engineer maintaining cloud infrastructure, availability, and performance for Hyland’s enterprise content intelligence platform. Automating deployments and supporting Kubernetes-based Cloud Services.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 24

Sezzle

201 - 500

💳 Fintech

👥 B2C

🛍️ Comércio Eletrônico

Senior Database Reliability Engineer scaling Sezzle’s interest-free installment-payment infrastructure. Improving reliability, observability, Kubernetes/AWS systems, and AI-enabled engineering productivity.

🇮🇳 Índia – Remoto

💵 $5.000 - $9.500 / mês

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 24

Akamai Technologies

5001 - 10000

🔒 Cibersegurança

Senior SRE optimizing Akamai’s distributed AI hardware infrastructure. Automating provisioning, observability, and incident response for reliable high-density cloud and edge systems.

🇮🇳 Índia – Remoto

💰 Post-IPO Equity em 2001-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório