Site Reliability Engineer

🕒 Agosto 25

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $120.000 - $140.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 1%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of System Automation Corporation

System Automation Corporation

51 - 200 funcionários

Fundada em 1973

💼 Consultoria

🏥 Saúde

⚖️ Jurídico

Consulting • Healthcare • Legal

A System Automation Corporation é uma empresa que oferece uma plataforma de aplicações SaaS de baixo código conhecida como Evoke. Esta plataforma apoia a transformação digital por meio de sistemas regulatórios empresariais adaptados para as necessidades de agências reguladoras e do setor público. O Evoke fornece gerenciamento completo de licenças e regulamentos, bem como capacidades de gerenciamento de registros, com o objetivo de modernizar operações e melhorar a produtividade de agências governamentais. Conhecida por sua rápida implantação e alta configurabilidade, a System Automation melhora a eficiência das operações regulatórias, permitindo soluções em nuvem seguras para o setor público.

Descrição

• Build, operate, and scale production systems in Microsoft Azure, including App Service, Networking, WAF, CosmosDB, and related infrastructure • Participate in an on-call rotation; triage, respond to, and resolve production incidents • Lead or contribute to blameless postmortems • Define and track SLOs/SLIs and error budgets with engineering and product teams • Design and maintain observability and APM tooling, including metrics, logs, tracing, dashboards, and alerting • Refine alert thresholds and runbooks to reduce noise and mean time to resolution • Reduce operational toil through scripting, self-healing systems, and repeatable processes • Build and maintain CI/CD pipelines enabling safe and frequent releases • Provision and manage infrastructure as code using Bicep • Follow change control and version control practices • Ensure application infrastructure meets security and compliance requirements such as SOC 2 and GovRAMP • Apply security best practices to infrastructure design and change management • Partner with the agile development team to translate business requirements into reliable technical solutions • Participate in technical design sessions • Produce clear documentation, including diagrams, runbooks, and architecture notes • Stay current on Azure capabilities, industry standards, and SRE best practices and bring recommendations to the team • Perform other duties as assigned

🎯 Requisitos

• 3+ years of experience in an IT Operations, DevOps, or SRE role • Hands-on technical experience with Microsoft Azure in a production environment • Experience with infrastructure as code — Terraform and/or Bicep • Proficiency in at least one scripting/programming language — Python or TypeScript • Experience working with REST and/or GraphQL APIs • Experience defining and tracking KPIs/SLOs for a web-based application • Comfortable participating in an on-call rotation • Solid understanding of networking fundamentals, HTTP/S, and observability principles • Ability to evaluate multiple technical approaches and recommend the most effective solution • Strong independent problem-solving skills balanced with effective collaboration in a team environment • Familiarity with software development lifecycle and programming/coding standards • Clear, professional communication, especially under incident pressure • Preferred: Experience with compliance audits (SOC 2 Type 2, GovRAMP) • Preferred: Familiarity with security frameworks (NIST, ISO 27001) • Preferred: AZ-104 certification, or equivalent Azure networking experience • Preferred: Experience with Node.js • Preferred: Experience with low-code platforms (Power Apps, Logic Apps) • Preferred: Familiarity with Scrum/Agile methodology and supporting tools (Confluence, JIRA, Git, Jenkins, Bamboo, TFS) • Preferred: Ability to translate business requirements directly into application/site behavior changes • Must be eligible to work in the U.S.

🏖️ Benefícios

• Eligible for the company's commission plan • Profit sharing • Conditional employment offer contingent upon successful completion of a drug screening and a fingerprint-based background investigation • Accommodation available during the application or selection process • Equal Opportunity Employer committed to an inclusive environment

Candidatar-se

Vagas Similares

🕒 Agosto 25

Blue River Technology

201 - 500

🌾 Agricultura

🤖 Inteligência Artificial

🔧 Hardware

Senior Site Reliability Engineer scaling Kubernetes platforms and cloud infrastructure. Supporting Blue River Technology’s autonomous robotics products through reliability, security, and observability.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 24

The Home Depot

10.000+ funcionários

🏗️ Construção

📦 Logística

🛒 Varejo

Senior Principal Reliability Engineer designing resilient infrastructure for Home Depot store systems, payments, and COM platforms. Guiding multiple engineering teams on reliability, cloud costs, and technology strategy.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $170.000 - $280.000 / ano

💰 Debt Financing em 2007-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 24

Symbotic

501 - 1000

🔧 Hardware

📦 Logística

🤖 Inteligência Artificial

Senior reliability manager scaling maintenance and asset performance across Exol’s automated warehouses. Driving uptime, safety, launches, vendor governance, and enterprise reliability standards.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 24

PingWind Inc. (SDVOSB)

51 - 200

💼 Consultoria

📦 Logística

🏥 Saúde

DevSecOps Engineer building and deploying secure cloud-based IAM systems for federal government clients. Maintaining highly available architectures, automated delivery, compliance, and infrastructure upgrades.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $93.000 - $128.000 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 24

Cisco

10.000+ funcionários

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Technical SRE leader keeping Splunk Cloud reliable for demanding enterprise customers. Owning critical incidents, customer stacks, automation strategy, and cloud infrastructure architecture.

🗣️🇺🇸🇬🇧 Inglês obrigatório