Site Reliability Engineer – II

🕒 Agosto 11

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟢 Júnior

🟡 Pleno

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 10%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Backblaze

Backblaze

201 - 500 funcionários

Fundada em 2007

🛍️ Comércio Eletrônico

🏢 Corporativo

💰 $5.000.000 Series A em 2012-07

Cloud Storage • eCommerce • Enterprise

A Backblaze é uma empresa de armazenamento em nuvem que oferece soluções de backup de dados escaláveis e seguras tanto para empresas quanto para indivíduos. Seu serviço B2 Cloud Storage oferece armazenamento de objetos compatível com S3, permitindo que os usuários protejam e gerenciem seus dados com preços transparentes. A Backblaze é especialista em serviços de backup automáticos e ilimitados para sistemas de computador, garantindo opções de proteção e recuperação de dados para os usuários, além de suportar a integração com aplicativos para funcionalidades aprimoradas.

Descrição

• Support the availability and durability of critical services across production environments • Monitor service health using SLIs, SLOs, and error budgets, escalating issues when thresholds are at risk • Participate in on-call rotations, incident response, and post-incident reviews • Follow ITIL/OSS processes for incident, change, problem, and capacity management • Develop automation for common operational tasks and reduce manual intervention • Contribute to monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Catchpoint, and ELK • Work with CI/CD pipelines, configuration management, and infrastructure-as-code tools including Terraform, Ansible, and Jenkins • Write Bash, Python, or Go scripts to improve system reliability and efficiency • Partner with engineering, product, and operations teams on resilient system design and operations • Assist with capacity planning and disaster recovery exercises • Work with vendors and service providers to troubleshoot issues and track SLA performance • Document systems, share learnings, and contribute to a reliability-minded engineering culture • Contribute to playbooks, runbooks, and operational documentation • Identify recurring issues and propose long-term improvements • Promote reliability-focused practices within development and operations teams

🎯 Requisitos

• Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent experience • 2–4 years of experience in site reliability, systems engineering, or operations • Exposure to large-scale, production-grade systems • Solid Linux systems administration and troubleshooting skills • Familiarity with monitoring, alerting, incident response, and root cause analysis • Proficiency in at least one scripting language: Python, Bash, or Go • Understanding of containers, including Kubernetes and Docker, and microservices concepts • Knowledge of incident response and operational best practices • Experience in a SaaS, service provider, or distributed systems environment (preferred) • Familiarity with ITIL/OSS practices and SLO/SLAs (preferred) • Experience with cloud platforms such as AWS, GCP, or Azure (preferred) • Ability to work independently, take ownership, and drive projects from problem discovery through resolution (preferred)

Candidatar-se

Vagas Similares

🕒 Agosto 5

Signalmash

51 - 200

💼 Consultoria

📦 Logística

🏥 Saúde

DevOps Engineer owning Kubernetes, CI/CD, PostgreSQL, observability, and security for Signalmash’s cloud communications platform. Improving reliability, deployment speed, recovery, and infrastructure costs from India.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

Neo4j

501 - 1000

☁️ SaaS

🤖 Inteligência Artificial

🏢 Corporativo

Cloud Operations Engineer managing and troubleshooting customer Neo4j database infrastructure. Supporting deployments, monitoring, upgrades, and incidents across AWS, Azure, Google Cloud, virtual, and bare-metal environments.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 29

Outmarket AI

11 - 50

🤖 Inteligência Artificial

🛡️ Seguros

☁️ SaaS

DevOps Engineer managing infrastructure and delivery platform for AI products. Focus on security, reliability, and observability while working in an AI-first environment.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 29

Granicus

501 - 1000

🏛️ Governo

☁️ SaaS

📋 Conformidade

Site Reliability Engineer 3 modernizing reliability engineering for Granicus with a focus on AIOps and automation. Improve service reliability and build scalable, resilient platforms for various workloads.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

fal

51 - 200

🤖 Inteligência Artificial

🔌 API

☁️ SaaS

Machine Learning Engineer focusing on the reliability and security of generative media model APIs at fal. Working with cutting-edge models and infrastructure in a remote setting.

🗣️🇺🇸🇬🇧 Inglês obrigatório