Senior Site Reliability Engineer

🕒 Agosto 4

🍂 Massachusetts – Remoto

infoinfo

💵 $121.400 - $218.600 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 2%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 funcionários

🔒 Cibersegurança

💰 Post-IPO Equity em 2001-07

Cloud Computing • Cybersecurity • Content Delivery

A Akamai Technologies é um provedor líder de serviços em nuvem especializado em oferecer soluções de segurança, computação em nuvem e entrega de conteúdo. A empresa oferece uma gama de serviços, como segurança de API, proteção contra DDoS e otimização de desempenho para aplicativos web, garantindo experiências de usuário seguras e confiáveis. Com uma infraestrutura global robusta, a Akamai capacita empresas a otimizar sua presença digital, protegendo contra diversas ameaças cibernéticas e melhorando o desempenho de aplicativos.

Descrição

• Oversee, scale, and optimize next-generation dedicated AI hardware infrastructure • Ensure uptime and reliability of AI hardware infrastructure offerings • Collaborate with product teams from early development through deployment to ensure reliability, scalability, and performance • Define key performance indicators and defend them when breached • Develop Python programmatic tooling and infrastructure-as-code utilities to automate fleet-wide provisioning and reduce operational toil • Integrate automated workflows across corporate ticketing systems for hardware and network break-fix events • Use AI utilities and LLM-assisted development to accelerate technical execution and system analysis • Improve availability, latency, and systemic health of high-density hardware environments using private cloud and compute technologies • Design telemetry pipelines, Prometheus/Grafana dashboards, and AI-based anomaly detection for bare-metal and virtualized environments • Participate in 24x7x365 on-call rotations and lead real-time incident management • Manage high-severity service disruption protocols through PagerDuty and Slack workflows • Partner with third-party infrastructure vendors and coordinate on-site field technicians

🎯 Requisitos

• 5 years of relevant experience • Bachelor's degree in Computer Engineering, Computer Science or equivalent • Tooling and coding ability in Python for scalable operational tools, API integrations, and automation frameworks • Hands-on experience with Prometheus, Grafana, OpenTelemetry, and Loki • Working understanding of advanced networking topologies and high-bandwidth routing/switching infrastructure • Knowledge of BGP and dual-stack IPv4/IPv6 networks • Experience designing new service rollouts, including operational readiness criteria, telemetry baselines, and alerting thresholds • Extensive experience building technical runbooks, leading complex incident response bridges, and driving blameless post-mortems • Ability to own ambiguous technical problems, coordinate cross-functional teams, and deliver production-grade solutions • Participation in 24x7x365 on-call rotations

🏖️ Benefícios

• Flexible work through Akamai's FlexBase program: at home, in an office, or a combination of both • Annual bonus or incentives • Equity awards • Employee Stock Purchase Plan (ESPP) • Healthcare • 401K savings plan • Company holidays • Vacation/PTO • Sick time • Parental leave • Employee assistance program • Mental and financial wellness support

Candidatar-se

Vagas Similares

🕒 Agosto 4

PrizePicks

201 - 500

🎮 Jogos

⚽ Esportes

Senior SRE ensuring reliable, scalable infrastructure for PrizePicks’ daily fantasy sports platform. Leading incident response, observability, and production systems engineering.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $120.000 - $175.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

Veeam Software

1001 - 5000

💼 Consultoria

📦 Logística

☁️ SaaS

Senior SRE building reliability engineering for Veeam Data Cloud’s Government and Sovereign Cloud SaaS platform. Designing Azure infrastructure, observability, incident response, and compliance-ready delivery practices.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $158.400 - $294.100 / ano

💰 $500.000.000 Private Equity Round em 2019-01

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

Syniti

1001 - 5000

🤝 B2B

🏢 Corporativo

Senior SRE automating Azure and AWS infrastructure for Syniti’s enterprise data platform. Supporting Kubernetes, CI/CD, observability, security, and compliance across global SaaS workloads.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $134.941 - $171.411 / ano

💰 Private Equity Round em 2017-08

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

The Access Group

5001 - 10000

💼 Consultoria

🏥 Saúde

🏨 Hospitalidade

Senior Site Reliability Engineer owning Azure, Kubernetes, Terraform, and production reliability for Access Group’s business management software platforms. Leading incident remediation, observability, compliance, and infrastructure architecture.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $165.000 - $185.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

Net Health

501 - 1000

🏥 Saúde

☁️ SaaS

🤖 Inteligência Artificial

DevOps Engineer designing secure AWS platforms and CI/CD automation for Net Health’s healthcare SaaS products. Owning cloud architecture, database performance, observability, security, and cost optimization.

🗣️🇺🇸🇬🇧 Inglês obrigatório