Senior DevOps – Platform Reliability Engineer

🕒 Maio 8

🗽 New York – Remoto

info

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Zingtree

Zingtree

11 - 50 funcionários

🏥 Saúde

🛡️ Seguros

📦 Logística

💰 $15.000.000 Series A em 2022-01

Healthcare • Insurance • Logistics

Zingtree é uma empresa que potencializa o suporte ao cliente por meio da automação de processos com IA, ajudando empresas a otimizar e simplificar processos complexos de suporte. Oferece ferramentas para criar, gerenciar e automatizar fluxos de trabalho de suporte, facilitando para que agentes de atendimento ao cliente e clientes resolvam questões de maneira eficiente. A Zingtree se integra a diversos sistemas de CRM e apoia múltiplas indústrias, incluindo centros de contato, saúde, seguros e serviços domiciliares. Através de seus fluxos de trabalho dinâmicos, do Author Assist AI e da automação de conformidade, ajuda empresas a melhorarem a experiência do cliente com tempos de resolução mais rápidos e controles de conformidade aprimorados.

Descrição

• Own and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments. • Automate infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform and CloudFormation. • Operate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers. • Manage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls. • Operate the data and event tier, including Aurora MySQL, ElastiCache/Redis, S3, and MSK (Kafka), with responsibility for backups, point-in-time recovery (PITR), and multi-AZ disaster recovery aligned to defined RTO/RPO objectives. • Build and maintain Lambda workloads where event-driven or serverless architectures are the right fit. • Build observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking. • Strengthen our security and compliance posture for SOC 2 and HIPAA, including least-privilege IAM, SCPs, secrets management, SAST/DAST, dependency and container scanning, image signing, AWS Config, Security Hub, GuardDuty, Inspector, and evidence automation. • Drive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls. • Build and evolve our AI-native DevOps capabilities.

🎯 Requisitos

• 5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS. • Strong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI. • Hands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades. • Strong AWS networking experience, including multi-account VPC design, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers. • Deep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication. • Experience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management. • Strong observability experience using Prometheus, Grafana, and OpenTelemetry. • Experience operating Argo CD at scale. • Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible. • Experience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront. • Experience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries. • Experience with Lambda and event-driven architectures. • Comfortable working with Python, Bash, and Linux systems. • Strong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security. • Familiarity with vulnerability scanning and compliance tooling.

🏖️ Benefícios

• Competitive compensation packages • Comprehensive health benefits: • 100% of employee premiums covered • 75%–80% of dependent premiums covered for most health, dental, and vision plans • 401(k) plans to support retirement planning (no employer matching currently) • Paid parental leave • Unlimited PTO • Flexible remote work from anywhere • Up to $200/month co-working reimbursement • Home office stipend: • Up to $500 for home office setup • $100/month for internet, phone, and related expenses

Candidatar-se

Vagas Similares

🕒 Maio 8

Flywire

1001 - 5000

🏥 Saúde

✈️ Turismo

📦 Logística

Manager II, Site Reliability Engineering at Flywire driving infrastructure reliability and performance. Leading SRE teams for global cloud-based systems and initiatives for production excellence.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $200.000 / ano

💰 $60.000.000 Series F em 2021-03

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 5

NextGen IT Services

51 - 200

🤝 B2B

🏢 Corporativo

🎯 Recrutamento

DevOps Engineer at NextGen IT Services responsible for building/maintaining CI/CD pipelines and cloud infrastructure. Focusing on operational efficiencies and security controls implementation.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 5

Ad Hoc LLC

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

DevOps Engineer III at Ad Hoc creating products that transform government services. Collaborating with teams to improve DevOps processes and deliver software efficiently.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $115.000 - $125.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 5

Flywire

1001 - 5000

🏥 Saúde

✈️ Turismo

📦 Logística

Site Reliability Engineering Manager overseeing reliability, automation, and performance for Flywire’s cloud infrastructure. Leading SRE efforts, mentoring team members, and ensuring production excellence.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $200.000 / ano

💰 $60.000.000 Series F em 2021-03

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 4

Twingate

51 - 200

🔒 Cibersegurança

☁️ SaaS

Senior DevOps engineer managing cloud infrastructure and optimizing CI/CD pipelines for Twingate's remote access solution. Join a passionate team across Israel, Europe, and the US.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $170.000 - $200.000 / ano

💰 $42.000.000 Series B em 2022-04

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório