Senior DevOps Engineer

Provável vaga fantasma

🕒 Fevereiro 2

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 60%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Shuru

Shuru

51 - 200 funcionários

Fundada em 2021

🤖 Inteligência Artificial

🤝 B2B

🏢 Corporativo

Artificial Intelligence • B2B • Enterprise

A Shuru é uma empresa de consultoria em produtos, IA e tecnologia que faz parceria com empresas para oferecer consultoria estratégica, desenvolvimento de produtos e softwares personalizados ponta a ponta, além de extensão de equipes de engenharia sob medida. Suas equipes de engenharia nativas em IA constroem aplicações de IA escaláveis, engenharia de dados e análises, cloud/DevOps e integrações de API para modernizar sistemas e acelerar a entrega de produtos. A Shuru opera globalmente com um modelo remoto-primeiro e enfatiza alta responsabilidade, design thinking e resultados mensuráveis para clientes empresariais e startups.

Descrição

• Kubernetes platform engineering (EKS-first) ● Design, build, and operate production-grade Kubernetes clusters (multi-nodegroup, autoscaling, upgrades, cluster add-ons). • Implement intelligent autoscaling using real metrics (queue depth, consumer lag, service latency) via tools like KEDA/Karpenter. • Own AWS environments end-to-end (VPC, IAM, EKS/ECS/EC2, ALB/ELB, S3, Route53, CloudWatch, RDS, SQS, Lambda). • Build reproducible infrastructure using Terraform, with strong review + change management practices. • Implement backup/DR patterns (e.g., snapshots, retention, automation) and safe rollouts. • Design infrastructure for data-intensive workloads: high-throughput ingestion, batch processing, and real-time streaming. • Understand and operate distributed systems at scale — consensus, partitioning, replication, and failure modes. • Build and maintain infrastructure for data pipelines, vector databases. • Design for horizontal scalability, ensuring systems handle growing data volumes and user traffic gracefully. • Build/own monitoring + logging from scratch and make it actionable (Prometheus/Grafana, ELK/EFK, alerting). • Define/partner on SLI/SLOs and incident response practices; improve reliability with data-driven changes. • Establish performance testing and production-like load testing environments. • Continuously reduce AWS spend via right-sizing, Spot strategies, reserved capacity planning, and architecture improvements. • Partner with engineering teams to diagnose bottlenecks (db queries, caching, queueing) and propose scalable solutions. • Optimize infrastructure costs for data-heavy workloads (storage tiering, compute scheduling, GPU utilization). • Improve cloud and cluster security posture (IAM, network policies, secrets management, least privilege). • Support SOC2 readiness/execution (controls, evidence automation, operational hardening). • Implement access management patterns.

🎯 Requisitos

• 7+ years in DevOps / SRE / Cloud Infra roles operating production systems. • Deep hands-on experience with Kubernetes in production. • Strong AWS fundamentals across compute/networking/storage/identity, including VPC, IAM, EC2/EKS, ALB, S3, Route53, CloudWatch, RDS, SQS. • Proven ability to build infra using Terraform (and strong IaC practices). • Production-grade observability experience: Prometheus + Grafana, and centralized logging (ELK/EFK or similar). • Experience scaling product infrastructure — you've grown systems from thousands to millions of requests, and understand capacity planning, bottleneck identification, and scaling patterns. • Solid understanding of distributed systems concepts: CAP theorem, consistency models, partitioning strategies, distributed consensus, and failure handling. • Strong understanding of databases and performance fundamentals. • CI/CD experience building reliable pipelines (Jenkins/Spinnaker/GitHub Actions equivalents), with safe deployment strategies. • Scripting/automation ability in Python and/or Bash (Go is a plus).

🏖️ Benefícios

• Competitive salary and benefits package. • Opportunity to work with a team of experienced product and tech leaders. • A flexible work environment with remote working options. • Continuous learning and development opportunities. • Chance to make a significant impact on diverse and innovative projects.

Candidatar-se

Vagas Similares

🕒 Dezembro 5, 2025

VELAIO

51 - 200

💼 Consultoria

DevOps Engineer designing and managing Azure cloud solutions in a dynamic environment. Collaborating with development teams and automating software delivery pipelines using Azure DevOps tools.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇪🇸 Espanhol obrigatório

🕒 Novembro 25, 2025

DKSH Portugal, Unipessoal, Lda.

11 - 50

🍽️ Alimentos e Bebidas

📦 Logística

💼 Consultoria

Senior DevOps Engineer responsible for designing, building, and maintaining Azure-based CI/CD pipelines. Working in a global technology team to support enterprise-scale solutions.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Novembro 21, 2025

Typeface

11 - 50

🤖 Inteligência Artificial

🤝 B2B

Forward Deployment Engineer at Typeface translating business needs into scalable technical architectures and building AI-driven applications. Collaborating with customer success and product teams on innovative solutions.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $100.000.000 Series B em 2023-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Novembro 13, 2025

Vimeo

1001 - 5000

💼 Consultoria

📣 Marketing

📦 Logística

Site Reliability Engineer responsible for optimizing Vimeo's cloud infrastructure and performance. Collaborating to ensure high reliability and efficiency across distributed systems.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $130.000 - $178.750 / ano

💰 $3.005.700.000 Private Equity Round em 2021-01

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Novembro 11, 2025

Aalyria

51 - 200

📡 Telecomunicações

🏢 Corporativo

☁️ SaaS

Site Reliability Engineer responsible for building an observability platform for satellite network systems. Develop and implement metrics, logging, and tracing solutions with cloud-native tools.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $115.000 - $135.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório