Site Reliability Engineer II

🕒 Junho 23

🇦🇷 Argentina – Remoto

⏰ Tempo Integral

🟢 Júnior

🟡 Pleno

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Backblaze

Backblaze

201 - 500 funcionários

Fundada em 2007

🛍️ Comércio Eletrônico

🏢 Corporativo

💰 $5.000.000 Series A em 2012-07

Cloud Storage • eCommerce • Enterprise

A Backblaze é uma empresa de armazenamento em nuvem que oferece soluções de backup de dados escaláveis e seguras tanto para empresas quanto para indivíduos. Seu serviço B2 Cloud Storage oferece armazenamento de objetos compatível com S3, permitindo que os usuários protejam e gerenciem seus dados com preços transparentes. A Backblaze é especialista em serviços de backup automáticos e ilimitados para sistemas de computador, garantindo opções de proteção e recuperação de dados para os usuários, além de suportar a integração com aplicativos para funcionalidades aprimoradas.

Descrição

• Support the availability and durability of critical services across production environments. • Monitor service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk. • Participate in on-call rotations, incident response, and post-incident reviews to drive service improvements. • Follow established ITIL/OSS processes (incident, change, problem, and capacity management). • Develop automation for common operational tasks, reducing manual intervention and toil. • Contribute to monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint,ELK). • Work with CI/CD pipelines, configuration management, and infrastructure as code tools (Terraform, Ansible, Jenkins). • Write scripts (Bash, Python, Go, etc.) to improve system reliability and efficiency. • Partner with engineering, product, and operations teams to support resilient system design and operations. • Assist in capacity planning and disaster recovery exercises. • Work with vendors and service providers to troubleshoot service issues and track SLA performance. • Document systems, share learnings, and help grow a reliability-minded engineering culture. • Contribute to playbooks, runbooks, and operational documentation. • Identify recurring issues and propose long-term improvements. • Promote reliability-focused practices within development and operations teams.

🎯 Requisitos

• Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience). • 2–4 years of experience in site reliability, systems engineering, or operations. • Exposure to large-scale, production-grade systems. • Solid Linux systems administration and troubleshooting skills. • Familiarity with service reliability concepts - monitoring, alerting, incident response, and root cause analysis. • Proficiency in at least one scripting language (Python, Bash, or Go). • Understanding of containers (Kubernetes, Docker) and microservices concepts. • Knowledge of incident response and operational best practices. • Strong problem-solving skills and willingness to learn new technologies. • Experience in a SaaS, service provider, or distributed systems environment is preferred.

🏖️ Benefícios

• We value being fair and good to our customers, partners, and employees. • Diversity, equity, and inclusion are at the core of our values. • We are committed to fostering a workforce where all employees feel a sense of belonging regardless of race, ethnicity, nationality, gender, sexual orientation, age, religion, socio-economic status, ability, veteran status, and education. • We believe that our dedication to cultivating a diverse workspace allows us to better serve our customers.

Candidatar-se

Vagas Similares

🕒 Junho 17

Pragmatike

11 - 50

💼 Consultoria

📣 Marketing

🎯 Recrutamento

SRE / Network Engineer handling decentralized cloud infrastructure for a European deep-tech company. Focus on bare-metal automation and advanced networking systems.

🇦🇷 Argentina – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 9

Teladoc Health

5001 - 10000

🏥 Saúde

👥 B2C

☁️ SaaS

DevOps Engineer III developing infrastructure and automating delivery for Teladoc Health products. Collaborating with development and QA teams to enhance productivity and secure systems.

🇦🇷 Argentina – Remoto

💰 $80.000.000 Post-IPO Debt - Teladoc Health em 2016-07

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 30

EY

10.000+ funcionários

🏥 Saúde

⚖️ Jurídico

📣 Marketing

DevOps Engineer responsible for designing and managing Azure DevOps solutions in a leading consulting firm. Collaborating with teams to automate and improve processes for client technology transformations.

🇦🇷 Argentina – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇪🇸 Espanhol obrigatório

🕒 Abril 1

Particle41

51 - 200

💼 Consultoria

📣 Marketing

☁️ SaaS

DevOps Engineer at Particle41 enhancing software delivery through automation and collaboration with development and operations teams. Managing Azure cloud infrastructure and ensuring system reliability.

🇦🇷 Argentina – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 1

Particle41

51 - 200

💼 Consultoria

📣 Marketing

☁️ SaaS

DevOps Engineer aligning software development and IT operations for Particle41. Focus on automation and GCP infrastructure management to ensure reliability and scalability.

🇦🇷 Argentina – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório