Site Reliability Engineer

🕒 6 dias atrás

🏄 California – Remoto

info

💵 $138.900 - $231.400 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Veeam Software

Veeam Software

1001 - 5000 funcionários

Fundada em 2006

💼 Consultoria

📦 Logística

☁️ SaaS

💰 $500.000.000 Private Equity Round em 2019-01

Consulting • Logistics • SaaS

A Veeam Software é líder global em resiliência e proteção de dados, oferecendo software de proteção de dados autogerenciável para ambientes híbridos e multi-cloud. Sua Veeam Data Platform oferece soluções abrangentes para backup, recuperação e segurança de dados, com princípios de confiança zero e ferramentas impulsionadas por inteligência artificial para inteligência de dados. As ofertas da Veeam incluem serviços de backup e armazenamento seguros para plataformas como Microsoft 365, AWS e Google Cloud, suportando cargas de trabalho diversas, incluindo ambientes virtuais, físicos e SaaS. Com uma reputação de inovação e confiança do cliente, a Veeam atende a uma ampla gama de indústrias, garantindo resiliência de dados contra interrupções, como ataques de ransomware. Suas soluções permitem que as empresas alcancem liberdade de dados, armazenamento seguro e gestão eficiente, reforçando sua posição como um dos principais fornecedores de software de backup e recuperação empresarial em todo o mundo.

Descrição

• Support the Veeam Data Cloud SaaS platform’s Government and Sovereign Cloud environment • Map platform systems, workloads, dependencies, and risk areas • Work with subject-matter experts to fill knowledge gaps and build onboarding material • Write and maintain runbooks, architecture documentation, and operational guides • Design highly available and fault-tolerant infrastructure on Azure, including Azure Government • Define SLIs, SLOs, and error budgets • Lead incident response and blameless postmortems, converting incidents into improvements • Identify reliability risks and develop remediation plans within compliance constraints • Define observability instrumentation requirements and drive implementation • Establish alerting, telemetry, and monitoring standards • Build automation to reduce toil and support fleet management • Participate in on-call rotations • Work with infrastructure as code, CI/CD, deployment automation, and configuration management in air-gapped or compliance-restricted environments • Build and maintain testing, canary deployment, and release validation pipelines • Integrate chaos engineering and monitoring tools • Collaborate across product, platform, security, legal, compliance, and operations teams • Own reliability problems end-to-end and drive solutions • Mentor engineers and spread SRE practices across the organization

🎯 Requisitos

• 7+ years in Software Engineering, including 3+ years in SRE, Platform Engineering, or similar roles • Experience across multi-service platforms • Experience with Government or Sovereign Cloud, such as Azure Government or AWS GovCloud • Experience in regulated compliance environments, including FedRAMP, CMMC, IL2/IL4/IL5, PCI-DSS, SOX, HIPAA, or HITRUST • Strong experience building and running production services on cloud infrastructure; Azure preferred, including Azure Government • Ability to learn large, complex platforms quickly with limited guidance and restricted environment access • Ability to independently investigate systems and produce clear documentation, risk assessments, and improvement plans • Experience with programming in TypeScript/JavaScript, Go, Java, C#, or similar • Experience with monitoring and observability tools such as Prometheus, Grafana, OpenTelemetry, or ELK Stack • Experience with infrastructure as code, including Terraform, Terragrunt, or Pulumi • Experience with container orchestration, especially Kubernetes • Experience with CI/CD and GitOps tooling, including GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or Dagger • Strong understanding of distributed systems, networking, and cloud-native architecture • Clear written and verbal communication skills

🏖️ Benefícios

• Unlimited paid time off • 12 paid holidays, including 4 global VeeaMe Days for self-care • 24 paid volunteer hours annually through Veeam Cares • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents • Medical, dental, and vision coverage starting on the first day • Mental health support, therapy sessions, and digital wellness tools via the Employee Assistance Program • 401(k) retirement plan with company matching contributions • Fertility, adoption, and surrogacy support through Maven • AirVet: 24/7 virtual veterinary care at no cost • Legal services, identity protection, and supplemental health insurance options • Tax-advantaged spending accounts for healthcare, dependent care, and commuting • On-demand learning libraries, mentoring, workshops, and learning events including the annual Global Day of Learning • Competitive compensation and benefits

Candidatar-se

Vagas Similares

🕒 6 dias atrás

Karat

201 - 500

👥 RH Tech

🏢 Corporativo

☁️ SaaS

Senior Deployment Engineer helping Karat, a technical interviewing company, implement and optimize enterprise interview frameworks. Advising clients, analyzing hiring performance, and delivering executive training.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $116.875 - $148.187 / ano

💰 Funding Round em 2022-04

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

Cisco

10.000+ funcionários

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Site Reliability Engineer architecting Cisco’s developer platform and infrastructure for cloud application delivery. Consolidating legacy tools and improving scalable, resilient engineering workflows.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

Cisco

10.000+ funcionários

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Customer Reliability Engineer resolving escalated Cisco Hypershield incidents across Nexus Smart Switches and on-premises Kubernetes. Improving reliability through diagnostics, runbooks, tooling, and engineering fixes.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

Ad Hoc LLC

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

DevOps Engineer III building AWS infrastructure, Terraform modules, and GitLab CI/CD platforms for government digital services. Supporting secure container deployments, observability, incident response, and federal compliance.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $125.000 - $142.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

The Access Group

5001 - 10000

💼 Consultoria

🏥 Saúde

🏨 Hospitalidade

Senior Site Reliability Engineer operating secure, scalable Azure platforms for Access, a business management software provider. Leading incident remediation, Kubernetes, Terraform, observability, compliance, and DevOps improvements.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $165.000 - $185.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório