Senior Site Reliability Engineer

🕒 Julho 1

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 52%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Empower

Empower

10.000+ funcionários

💸 Finanças

💳 Fintech

👥 B2C

Finance • Fintech • B2C

A Empower é uma fornecedora líder de serviços financeiros focada em ajudar indivíduos e organizações a alcançar a liberdade financeira através do planejamento de aposentadoria e da gestão de investimentos. Atendendo a mais de 19 milhões de americanos, a Empower oferece um conjunto abrangente de serviços relacionados a finanças, incluindo planejamento inteligente e aconselhamento de investimento, além de ferramentas como o Empower Personal Dashboard™ para uma visão financeira completa. A empresa é renomada como uma das principais provedoras de planos de aposentadoria e trabalha de perto com investidores pessoais, poupadores de planos no local de trabalho, patrocinadores de planos e profissionais financeiros. A Empower também é reconhecida por suas iniciativas em Diversidade, Equidade, Inclusão e tem um compromisso social que fortalece o impacto na comunidade.

Descrição

• Design and implement highly available, fault-tolerant systems supporting critical financial transactions • Architect infrastructure solutions using AWS best practices, optimizing for cost, performance, and reliability • Lead complex incident response efforts and coordinate across teams to restore service rapidly • Drive postmortem processes for high-severity incidents and ensure action items are completed • Establish and track SLOs and SLIs for key services • Design and implement disaster recovery strategies and business continuity plans • Build Infrastructure as Code solutions using Terraform, including modules, workspaces, and state management • Architect and optimize multi-cluster EKS environments with pod autoscaling, cluster autoscaling, and resource optimization • Design observability strategies using Datadog and Splunk, including metrics, dashboards, and alerting • Implement progressive delivery mechanisms such as canary and blue-green deployments within GitOps workflows • Build automation frameworks to reduce operational toil and improve team efficiency • Partner with development teams on application reliability, design reviews, and architectural guidance • Mentor junior and intermediate SREs, conduct code reviews, and provide technical coaching • Contribute to architectural decisions affecting platform reliability and scalability • Evangelize SRE best practices across the engineering organization • Participate in on-call rotations and reduce on-call burden • Implement and maintain zero-trust security controls across infrastructure • Ensure systems meet financial services regulatory requirements and internal compliance standards • Conduct security reviews of infrastructure changes and deployment processes • Participate in audit preparations and respond to compliance-related inquiries

🎯 Requisitos

• Bachelor's degree in Computer Science, Information Systems or similar emphasis, or equivalent experience • 4-7 years of experience in Site Reliability Engineering (or equivalent), with a track record of operating large-scale production systems • Deep expertise in AWS, with hands-on experience across a broad range of services and architectural patterns • Advanced Kubernetes knowledge, including custom resources, operators, and cluster federation concepts • Expert-level proficiency in Terraform, including module development, state management, and complex workflow orchestration • Strong programming skills in Python and/or Go, with ability to develop production-quality tools and services • Production experience implementing observability at scale using Datadog, Splunk, or similar platforms • Demonstrated experience establishing and maintaining CI/CD pipelines at enterprise scale • Deep understanding of GitOps principles and experience with tools like ArgoCD or Flux • Proven ability to lead complex incident response and conduct thorough postmortems • Strong understanding of networking, security, and infrastructure design patterns • Experience mentoring engineers and conducting technical training • Preferred: Experience in financial services or payments industry • Preferred: Deep knowledge of compliance frameworks (SOC 2, PCI DSS, FINRA) • Preferred: AWS certifications (Solutions Architect Professional, DevOps Engineer Professional) • Preferred: CKA and/or CKAD certifications • Preferred: Experience with service mesh implementations (Istio, Linkerd, Consul) • Preferred: Background in chaos engineering and fault injection testing • Preferred: Experience with FinOps and cloud cost optimization • Preferred: Contributions to open-source projects in the SRE/DevOps space • Preferred: Experience implementing Operational Excellence strategies

🏖️ Benefícios

• Flexible work environment • Fluid career paths • Internal mobility opportunities • Well-being support • Work-life balance • Inclusive and welcoming work environment • Volunteering opportunities

Candidatar-se

Vagas Similares

🕒 Junho 23

SigNoz

11 - 50

☁️ SaaS

🏢 Corporativo

SRE responsible for the reliability and operability of SigNoz cloud platform while scaling observability systems and ingest pipelines. Work in a fast-paced, remote-first environment with a high-caliber team.

🇮🇳 Índia – Remoto

💵 ₹5.000.000 - ₹10.000.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 19

BETSOL

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior Cloud Engineer at BETSOL building and operating cloud portal workloads across Azure and GCP. Focused on DevOps and DevSecOps with AI-first development practices.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 8

RevMind srl

11 - 50

🤖 Inteligência Artificial

💊 Farmacêutico

🏢 Corporativo

Senior Cloud DevOps Engineer owning secure, scalable AWS infrastructure for Revmind Labs AI's enterprise AI and analytics systems. Automating deployments, observability, security, and production reliability.

🇮🇳 Índia – Remoto

💵 ₹3.000.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 1

OpenAI

201 - 500

🤖 Inteligência Artificial

☁️ SaaS

🏢 Corporativo

Partner AI Deployment Engineer responsible for AWS deployment strategies and technical leadership in OpenAI. Guiding enterprise customers from ideation to production while influencing joint account strategy.

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 21

Akamai Technologies

5001 - 10000

🔒 Cibersegurança

🏢 Corporativo

📱 Mídia

Senior Site Reliability Engineer focusing on developing solutions for automation and efficiency with Akamai's Compute products. Enhance reliability and operational excellence in customer-facing applications and infrastructure.

🗣️🇺🇸🇬🇧 Inglês obrigatório