Site Reliability Engineer II

🕒 Julho 15

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟢 Júnior

🟡 Pleno

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Backblaze

Backblaze

201 - 500 funcionários

Fundada em 2007

🛍️ Comércio Eletrônico

🏢 Corporativo

💰 $5.000.000 Series A em 2012-07

Cloud Storage • eCommerce • Enterprise

A Backblaze é uma empresa de armazenamento em nuvem que oferece soluções de backup de dados escaláveis e seguras tanto para empresas quanto para indivíduos. Seu serviço B2 Cloud Storage oferece armazenamento de objetos compatível com S3, permitindo que os usuários protejam e gerenciem seus dados com preços transparentes. A Backblaze é especialista em serviços de backup automáticos e ilimitados para sistemas de computador, garantindo opções de proteção e recuperação de dados para os usuários, além de suportar a integração com aplicativos para funcionalidades aprimoradas.

Descrição

• Support the availability and durability of critical services across production environments. • Monitor service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk. • Participate in on-call rotations, incident response, and post-incident reviews to drive service improvements. • Follow established ITIL/OSS processes (incident, change, problem, and capacity management). • Develop automation for common operational tasks, reducing manual intervention and toil. • Contribute to monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint, ELK). • Work with CI/CD pipelines, configuration management, and infrastructure as code tools (Terraform, Ansible, Jenkins). • Write scripts (Bash, Python, Go, etc.) to improve system reliability and efficiency. • Partner with engineering, product, and operations teams to support resilient system design and operations. • Assist in capacity planning and disaster recovery exercises. • Work with vendors and service providers to troubleshoot service issues and track SLA performance. • Document systems, share learnings, and help grow a reliability-minded engineering culture. • Contribute to playbooks, runbooks, and operational documentation. • Identify recurring issues and propose long-term improvements. • Promote reliability-focused practices within development and operations teams.

🎯 Requisitos

• Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience). • 2–4 years of experience in site reliability, systems engineering, or operations. • Exposure to large-scale, production-grade systems. • Solid Linux systems administration and troubleshooting skills. • Familiarity with service reliability concepts - monitoring, alerting, incident response, and root cause analysis. • Proficiency in at least one scripting language (Python, Bash, or Go). • Understanding of containers (Kubernetes, Docker) and microservices concepts. • Knowledge of incident response and operational best practices.

🏖️ Benefícios

• Competitive salary • Flexible working hours • Professional development budget • Home office setup allowance • Global team events

Candidatar-se

Vagas Similares

🕒 Julho 14

Raya

51 - 200

🌍 Impacto Social

👥 B2C

📱 Mídia

DevSecOps Engineer improving AWS/EKS security and driving collaboration between DevOps and engineering teams at Raya. Focused on hardening systems and closing security findings across the platform.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 14

General Dynamics Information Technology

10.000+ funcionários

💼 Consultoria

🏥 Saúde

📦 Logística

Deploy Automation Engineer at GDIT developing automated deployment pipelines for healthcare organizations. Collaborating with teams and ensuring continuous integration and delivery for various applications.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 14

Nymbus

201 - 500

💼 Consultoria

📣 Marketing

🏦 Bancário

Release Engineer managing production deployments and stability on Nymbus platform. Partnering with Release Coordination to ensure successful deployments in a remote-first environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $115.000 - $125.000 / ano

⏰ Tempo Integral

🟢 Júnior

🟡 Pleno

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🚫👨‍🎓 Sem graduação necessária

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 14

SambaNova Systems

201 - 500

🤖 Inteligência Artificial

🔧 Hardware

🏢 Corporativo

Forward Deployment Engineer embedding with enterprise customers to design and deploy GenAI applications. Collaborating across strategic product offerings to drive value implementation.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 14

11:11 Systems

1001 - 5000

🤝 B2B

🔒 Cibersegurança

🏢 Corporativo

Infrastructure Deployment Engineer at 11:11 Systems leading deployments of hardware infrastructure across global data centers. Coordinating cross-functional teams, optimizing workflows, and ensuring project execution.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $94.000 - $130.500 / ano

💰 Private equity em 2021-10

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório