Staff Site Reliability Engineer – Volcano

🕒 Junho 22

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $210.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Kong Inc.

Kong Inc.

201 - 500 funcionários

Fundada em 2017

💼 Consultoria

📦 Logística

🔌 API

💰 $100.000.000 Series D em 2021-02

Consulting • Logistics • API

Kong Inc. é uma empresa que oferece uma plataforma de APIs abrangente, projetada para facilitar o gerenciamento de APIs, a integração de IA e a produtividade de desenvolvedores. A companhia disponibiliza soluções como Kong Gateway, Kong Konnect e uma variedade de outras ferramentas voltadas ao gerenciamento e à otimização do ciclo de vida de APIs. A plataforma da Kong oferece suporte a ambientes multicloud e foi criada para entregar alto desempenho e segurança. É reconhecida pela Gartner como líder em gerenciamento de APIs e sustenta inovações em setores como serviços financeiros, saúde e tecnologia. A empresa enfatiza flexibilidade, segurança e velocidade, tornando-se a escolha preferida de organizações que buscam aprimorar seus serviços digitais por meio de APIs. A Kong também apoia uma comunidade robusta de desenvolvedores e oferece um amplo conjunto de integrações e plugins para simplificar o gerenciamento e as operações de APIs.

Descrição

• Own reliability for Volcano end-to-end: Define and drive SLOs, error budgets, and incident response practices for all Volcano services — edge deployments, managed Postgres, auth, realtime, storage, and the control plane. • Architect the platform's infrastructure: Design and build the multi-region Kubernetes infrastructure, networking, and data plane that powers Volcano's edge deployment pipeline and backend-as-a-service capabilities. • Build the GitOps and CI/CD backbone: Establish deployment automation, canary pipelines, and preview environment provisioning using ArgoCD, Helm, and Terraform/Terragrunt — setting patterns the broader team will follow. • Scale managed data services: Design, operate, and harden multi-tenant PostgreSQL clusters, Redis caching layers, and object storage — with a focus on data isolation, performance, and disaster recovery. • Drive observability from day one: Instrument every Volcano service with meaningful SLIs; build dashboards, alerts, and runbooks using Datadog, Prometheus, and Grafana before services go live, not after incidents. • Lead cross-functional reliability work: Collaborate with the OCTO team, product engineering, and security to bake reliability and compliance into Volcano's architecture — not bolt it on later. • Set SRE culture and standards: Mentor engineers across Volcano's contributing teams on reliability principles; lead postmortems, define on-call practices, and build a blameless engineering culture. • Evaluate and adopt emerging technologies: Given Volcano's greenfield nature, evaluate and make architectural decisions on edge runtimes, serverless compute, vector databases, and AI-native infrastructure components.

🎯 Requisitos

• BS in Computer Science or equivalent; substantial experience at Staff or Principal IC level in SRE/Platform Engineering. • Proven track record building SRE or platform engineering practices for developer-facing platforms or PaaS/SaaS products — ideally at greenfield stage. • Deep Kubernetes expertise: multi-tenant cluster design, networking (CNI, service mesh, ingress), autoscaling, and security hardening.

🏖️ Benefícios

• healthcare benefits • 401(k) plan • short and long term disability benefits • basic life and AD&D insurance

Candidatar-se

Vagas Similares

🕒 Junho 20

Gorilla Logic

501 - 1000

💼 Consultoria

📣 Marketing

📦 Logística

Technical Engineering Manager leading high-performing cloud and DevOps teams. Guiding architecture and delivery of scalable, reliable, and secure cloud solutions for clients.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 17

ClassWallet

11 - 50

💳 Fintech

📚 Educação

🏛️ Governo

DevOps Engineer optimizing cloud infrastructure and deployment pipelines for fintech company. Redefining public funds management and ensuring system reliability with high compliance standards.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $500.000 Debt Financing em 2020-05

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 16

Domino Data Lab

201 - 500

🤖 Inteligência Artificial

🏢 Corporativo

☁️ SaaS

Staff Site Reliability Engineer working on AI-assisted reliability tooling at Domino Data Lab. Leading incident response and enhancing system observability for critical services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $200.000 - $230.000 / ano

💰 Series F em 2022-06

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 15

TrueML

51 - 200

💼 Consultoria

🏥 Saúde

📣 Marketing

Sr. Security Engineer leading integration of security across the software development lifecycle at TrueML. Engaging in security automation, cloud security, and innovative AI solutions.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $122.090 - $160.000 / mês

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 15

Stord

1001 - 5000

🍽️ Alimentos e Bebidas

💼 Consultoria

📣 Marketing

Staff Site Reliability Engineer focusing on security to enhance GCP and CI/CD processes. Join Stord in advancing tech solutions for better consumer experiences.

🗣️🇺🇸🇬🇧 Inglês obrigatório