Senior Site Reliability Engineer

🕒 Julho 2

🌐 Estados Unidos, Canadá – Remoto

info

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Akuity

Akuity

11 - 50 funcionários

Fundada em 2021

🏢 Corporativo

☁️ SaaS

Enterprise • SaaS • Cloud

Akuity é uma plataforma projetada para equipes de engenharia de plataforma e DevOps, especializada em soluções de implantação Kubernetes. Ela melhora a eficiência da engenharia ao aproveitar os princípios do GitOps para facilitar a entrega contínua, a orquestração de contêineres e a entrega progressiva. A plataforma Akuity oferece recursos como medidas de segurança robustas, escalabilidade e monitoramento em tempo real, tornando-a uma escolha confiável para empresas líderes que gerenciam o Argo CD em grande escala.

Descrição

• Own SLI/SLO/SLA definitions for the Akuity SaaS platform and drive continuous improvement against them • Design, instrument, and maintain observability systems (metrics, logs, traces) across multi-region AWS infrastructure • Identify reliability gaps, lead blameless post-mortems, and close the loop with permanent fixes • Partner with engineering teams to build reliability into new features before they ship to production • Participate in an on-call rotation and act as incident commander for high-severity production events • Build and maintain runbooks, escalation paths, and incident playbooks that keep mean time to resolution low • Drive improvements to alerting fidelity; reduce noise, increase signal, eliminate toil • Lead post-incident reviews with clear timelines, root cause analysis, and follow-through on action items

🎯 Requisitos

• 5+ years of SRE, platform engineering, or production operations experience in a SaaS environment • Deep hands-on Kubernetes expertise; you understand the scheduler, networking, storage, and autoscaling at a level where you can debug anything • Strong AWS fundamentals across compute (EC2, EKS), networking (VPC, NLB, Route53), storage (S3, RDS), and IAM • Experience defining and operating against SLOs in production; you've written error budgets, not just read about them • Proficiency with observability tooling (Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent) • Solid scripting and automation skills; Go, Python, Bash, or similar; you automate what you touch • Strong written communication: clear runbooks, sharp incident reports, thoughtful post-mortems • Live within US time zones (Pacific through Eastern), including Canada and other regions

🏖️ Benefícios

• Competitive compensation, commensurate with experience • Equity participation in a well-funded, growing company • Fully remote: work from anywhere within US time zones (Pacific through Eastern), including Canada and other regions • Home office stipend and equipment budget • Flexible time off and a culture that respects it • Work directly with the engineers who built Argo CD and Kargo; you'll learn a lot here • US-based employees receive full benefits, including comprehensive health, dental, and vision coverage. Candidates based outside the US will be engaged as contractors.

Candidatar-se

Vagas Similares

🕒 Julho 2

Sanity.io

51 - 200

💼 Consultoria

📣 Marketing

📦 Logística

SRE managing scalable content operations infrastructure for AI-powered platform. Collaborating with dev teams and ensuring reliability for high request volume systems.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 2

Global Alliant Inc

51 - 200

💼 Consultoria

🏥 Saúde

📦 Logística

Senior Full Stack Software Engineer in Agile teams supporting federal technology initiatives. Responsible for building secure, scalable, cloud-native applications using modern tech stacks.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 1

Upstart

1001 - 5000

🚘 Automotivo

💼 Consultoria

🏥 Saúde

DevOps Engineer focused on enhancing cloud infrastructure reliability and performance at Upstart. Partnering with cross-functional teams to optimize Kubernetes and AWS cloud services.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 1

Buyers Edge Platform

501 - 1000

📦 Logística

🏭 Manufatura

🏨 Hospitalidade

DevOps Engineer improving reliability and operational efficiency of hosted infrastructure and applications for Buyers Edge Platform. Designing and building tooling to streamline workflows in foodservice tech.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $425.000.000 Private Equity Round em 2024-04

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 1

Ad Hoc LLC

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior DevSecOps Engineer working with federal enterprise cloud platform. Designing secure, automated CI/CD pipelines and collaborating with government stakeholders.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $145.000 - $160.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório