Senior Manager, DevOps

Vaga não está no LinkedIn

🕒 Abril 21

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of TrueML

TrueML

51 - 200 funcionários

💼 Consultoria

🏥 Saúde

📣 Marketing

Consulting • Healthcare • Marketing

A TrueML é uma empresa líder no setor de fintech, conhecida por suas soluções inovadoras que priorizam a experiência do cliente na indústria de serviços financeiros. A empresa, juntamente com sua família de companhias como a TrueAccord, foca no desenvolvimento de plataformas de comunicação e produtos inteligentes, digitais desde o início, que revolucionam a experiência do consumidor na gestão de saúde financeira. A TrueML aproveita a expertise de uma equipe dinâmica de cientistas de dados, especialistas em serviços financeiros e especialistas em experiência do cliente para criar tecnologia que resolve obstáculos ao bem-estar financeiro dos consumidores, garantindo inclusão e acessibilidade nos sistemas financeiros. Fundada em 2013 por Ohad Samet, a TrueML continua a transformar os serviços financeiros tradicionais, tornando-os mais amigáveis e eficazes para o consumidor.

Descrição

• Define and execute the long-term strategic vision for Infrastructure as Code (IaC), CI/CD evolution, and cloud-native architecture to support TrueML’s scaling needs. • Lead the design and implementation of self-service internal platforms to reduce developer cognitive load, enabling feature teams to deploy and manage services with minimal friction at increased velocity. • Act as the primary stakeholder for cloud spend (AWS); drive cost-optimization initiatives and lead contract negotiations for the DevOps toolstack and third-party vendors. • Ensure the infrastructure architecture supports strict High Availability (HA) requirements and robust Disaster Recovery (DR) protocols, maintaining system integrity across multiple regions. • Oversee the implementation and evolution of comprehensive monitoring, logging, and distributed tracing systems, leveraging AIOps to move from reactive to predictive system maintenance. • Champion security by design by integrating automated vulnerability scanning, secret management, and compliance checks directly into the automated build pipelines. • Serve as the ultimate escalation point for major production outages, facilitating blameless post-mortem reviews that focus on systemic improvements rather than individual error. • Maintain deep technical currency in container orchestration (Kubernetes), serverless patterns, and modern automation frameworks to provide meaningful mentorship and architectural guidance to senior engineering staff.

🎯 Requisitos

• Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience. • 10+ years of experience in DevOps, Site Reliability Engineering (SRE), or Software Engineering; 5+ years of experience managing engineers • Expert-level mastery with AWS and experience managing multi-region, high-availability deployments • Advanced experience with Kubernetes (K8s) and Docker, including cluster management, networking, and scaling in a production environment. • Proficiency in Terraform to drive consistency and automation across all infrastructure layers. Experience with Atlantis is a plus. • Deep experience designing and maintaining complex pipelines (GitHub Actions, GitLab CI, or Jenkins) and mastery of scripting languages like Python, Go, or Bash. • Hands-on experience with modern monitoring, observability, and tracing stacks (Datadog, Observe) and a firm grasp of SRE principles (SLIs/SLOs/Error Budgets). • Experience acting as an Incident Commander for high-severity outages and fostering a "blameless" post-mortem culture. • Demonstrated ability to influence executive leadership and collaborate cross-functionally with Product, Engineering, and Security teams. • Experience integrating AI-assisted productivity tools (Cline, GitHub Copilot) into the engineering workflow to accelerate delivery.

Candidatar-se

Vagas Similares

🕒 Abril 21

Sweed POS

11 - 50

💼 Consultoria

📦 Logística

📣 Marketing

DevOps Engineer optimizing infrastructure and implementing automation for Sweed's cannabis retail platform. Collaborate with global teams to enhance development and deployment processes.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 20

URBN (Urban Outfitters, Anthropologie Group, Free People & Nuuly)

10.000+ funcionários

👥 B2C

🛒 Varejo

👗 Moda

Senior DevOps Engineer optimizing cloud infrastructure on GCP for Nuuly. Leading CI/CD initiatives and collaborating with developers to enhance system performance.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 20

Skydio

501 - 1000

🎖️ Defesa

🏭 Manufatura

📦 Logística

Deployment Engineer managing technical implementation and support for Skydio's cutting-edge cloud connected products. Collaborating across internal teams and directly with customers to ensure success.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $115.000 - $135.000 / ano

💰 $170.000.000 Series E - Skydio em 2024-11

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 17

Gifthealth

501 - 1000

🏥 Saúde

📦 Logística

💼 Consultoria

Lead Site Reliability Engineer at Gifthealth developing scalable Ruby on Rails applications. Responsible for embedding reliability, automation, and DevOps practices into software systems.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $123.000 - $154.000 / ano

💰 $40.000.000 Private Equity Round - GiftHealth em 2023-04

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 17

Quzara LLC

11 - 50

💼 Consultoria

🏥 Saúde

🎖️ Defesa

Site Reliability Engineer ensuring resilience and security of Azure Government environments supporting Quzara's Cybertorch platform. Focus on infrastructure engineering, compliance, and automation strategies.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório