Senior Site Reliability Engineer

🕒 Agosto 4

🏄 California – Remoto

infoinfo

💵 $158.400 - $294.100 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 9%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Veeam Software

Veeam Software

1001 - 5000 funcionários

Fundada em 2006

💼 Consultoria

📦 Logística

☁️ SaaS

💰 $500.000.000 Private Equity Round em 2019-01

Consulting • Logistics • SaaS

A Veeam Software é líder global em resiliência e proteção de dados, oferecendo software de proteção de dados autogerenciável para ambientes híbridos e multi-cloud. Sua Veeam Data Platform oferece soluções abrangentes para backup, recuperação e segurança de dados, com princípios de confiança zero e ferramentas impulsionadas por inteligência artificial para inteligência de dados. As ofertas da Veeam incluem serviços de backup e armazenamento seguros para plataformas como Microsoft 365, AWS e Google Cloud, suportando cargas de trabalho diversas, incluindo ambientes virtuais, físicos e SaaS. Com uma reputação de inovação e confiança do cliente, a Veeam atende a uma ampla gama de indústrias, garantindo resiliência de dados contra interrupções, como ataques de ransomware. Suas soluções permitem que as empresas alcancem liberdade de dados, armazenamento seguro e gestão eficiente, reforçando sua posição como um dos principais fornecedores de software de backup e recuperação empresarial em todo o mundo.

Descrição

• Build reliability practices for Veeam Data Cloud’s Government and Sovereign Cloud environment • Map platform systems, dependencies, workloads, and risk areas • Create onboarding materials, runbooks, architecture documentation, and operational guides • Design highly available and fault-tolerant infrastructure on Azure and Azure Government • Define SLIs, SLOs, and error budgets • Lead incident response and blameless postmortems • Identify reliability risks and develop remediation plans within compliance constraints • Define observability instrumentation requirements and drive implementation • Establish alerting, telemetry, and monitoring standards • Build automation to reduce toil and support fleet management • Participate in on-call rotations • Work with IaC, CI/CD, deployment automation, and configuration management in air-gapped or restricted environments • Build and maintain testing, canary deployment, and release validation pipelines • Integrate chaos engineering and monitoring tools • Collaborate with product, platform, security, legal, compliance, and operations teams • Own reliability problems end-to-end and drive solutions • Mentor engineers and spread SRE practices across Veeam

🎯 Requisitos

• 7+ years in Software Engineering, including 3+ years in SRE, Platform Engineering, or similar, across multi-service platforms • Experience with Government or Sovereign Cloud, such as Azure Government or AWS GovCloud • Experience in regulated compliance environments, including FedRAMP, CMMC, IL2/IL4/IL5, PCI-DSS, SOX, HIPAA, or HITRUST • Strong experience building and running production services on cloud infrastructure; Azure preferred, including Azure Government • Ability to learn large, complex platforms quickly with limited guidance and restricted environment access • Ability to investigate systems independently and produce clear documentation, risk assessments, and improvement plans • Programming skills in TypeScript/JS, Go, Java, C#, or similar • Experience with monitoring and observability tools such as Prometheus, Grafana, OpenTelemetry, or ELK stack • Experience with IaC tools such as Terraform, Terragrunt, or Pulumi • Experience with container orchestration using Kubernetes • Experience with CI/CD and GitOps tooling such as GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or Dagger • Strong understanding of distributed systems, networking, and cloud-native architecture • Clear written and verbal communication skills • Bonus: experience on B2B SaaS platforms in regulated or government markets • Bonus: background in chaos engineering, resilience testing, or performance/load testing • Bonus: experience building an SRE or reliability function from scratch • Bonus: experience across modern cloud-native and legacy systems • Bonus: familiarity with AI-first development workflows using LLM-powered tools

🏖️ Benefícios

• Unlimited paid time off • 12 paid holidays, including 4 global VeeaMe Days for self-care • 24 paid volunteer hours annually through Veeam Cares • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents • Medical, dental, and vision coverage starting on the first day • Mental health support, therapy sessions, and digital wellness tools via the Employee Assistance Program • 401(k) retirement plan with company matching contributions • Fertility, adoption, and surrogacy support through Maven • AirVet 24/7 virtual veterinary care at no cost • Legal services, identity protection, and supplemental health insurance options • Tax-advantaged spending accounts for healthcare, dependent care, and commuting • On-demand learning libraries including LinkedIn Learning and O’Reilly • Mentoring, workshops, and learning events including the annual Global Day of Learning • Competitive performance-based bonus included in total target compensation • Comprehensive benefits package including health coverage, retirement plans, and unlimited time off

Candidatar-se

Vagas Similares

🕒 Agosto 4

Syniti

1001 - 5000

🤝 B2B

🏢 Corporativo

Senior SRE automating Azure and AWS infrastructure for Syniti’s enterprise data platform. Supporting Kubernetes, CI/CD, observability, security, and compliance across global SaaS workloads.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $134.941 - $171.411 / ano

💰 Private Equity Round em 2017-08

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

The Access Group

5001 - 10000

💼 Consultoria

🏥 Saúde

🏨 Hospitalidade

Senior Site Reliability Engineer owning Azure, Kubernetes, Terraform, and production reliability for Access Group’s business management software platforms. Leading incident remediation, observability, compliance, and infrastructure architecture.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $165.000 - $185.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

Net Health

501 - 1000

🏥 Saúde

☁️ SaaS

🤖 Inteligência Artificial

DevOps Engineer designing secure AWS platforms and CI/CD automation for Net Health’s healthcare SaaS products. Owning cloud architecture, database performance, observability, security, and cost optimization.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 3

CLARA Analytics

51 - 200

💼 Consultoria

🏥 Saúde

⚖️ Jurídico

DevOps Engineer at CLARA Analytics improving infrastructure-as-code practices in a fully remote environment. Collaborating with cross-functional teams and automating workflows for an AI-powered analytics platform.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 1

JFrog

1001 - 5000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior DevOps Engineer helping JFrog customers build CI/CD platforms with JFrog, Docker, Kubernetes, and cloud technologies. Influencing product roadmaps and training DevOps communities.

🗣️🇺🇸🇬🇧 Inglês obrigatório