Site Reliability Engineer II, DBA

🔥 0 minutes ago

🌐 Argentina, Colombia, +2 more countries – Remote

infoinfo

⏰ Full Time

🟢 Junior

🟡 Mid-level

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Backblaze

Backblaze

201 - 500 employees

Founded 2007

🛍️ eCommerce

🏢 Enterprise

💰 $5M Series A on 2012-07

Cloud Storage • eCommerce • Enterprise

Backblaze is a cloud storage company that provides scalable and secure data backup solutions for both businesses and individuals. Their B2 Cloud Storage service offers S3 compatible object storage, allowing users to easily protect and manage their data with transparent pricing. Backblaze specializes in automatic and unlimited backup services for computer systems, ensuring data protection and recovery options for users, while also supporting integration with applications for enhanced functionality.

📋 Description

• Operate and maintain high-availability database systems, primarily Vitess (distributed MySQL) and Cassandra, using established architecture and runbooks • Optimize database performance through query tuning, indexing strategies, and schema design • Execute documented backup, recovery, and replication procedures and escalate architecture-level changes to senior DBA SREs • Ensure database security compliance and access control • Support availability and durability of critical services across production environments • Monitor service health using SLIs, SLOs, and error budgets, escalating issues when thresholds are at risk • Participate in on-call rotations, incident response, and post-incident reviews • Follow established ITIL/OSS processes for incident, change, problem, and capacity management • Develop automation for common operational tasks to reduce manual intervention and toil • Contribute to monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Catchpoint, and ELK, and integrate runbooks with FireHydrant • Work with CI/CD pipelines, configuration management, and infrastructure as code tools including Terraform, Ansible, and Jenkins • Write Bash, Python, or Go scripts to improve system reliability and efficiency • Partner with engineering, product, and operations teams on resilient system design and operations • Assist with capacity planning and disaster recovery exercises • Work with vendors and service providers to troubleshoot service issues and track SLA performance • Document systems, share learnings, and contribute to a reliability-minded engineering culture • Contribute to playbooks, runbooks, and operational documentation • Identify recurring issues and propose long-term improvements • Promote reliability-focused practices within development and operations teams

🎯 Requirements

• Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience) • 2–4 years of experience in site reliability, systems engineering, or operations centered around database systems • Exposure to large-scale, production-grade systems • Solid Linux systems administration and troubleshooting skills • Familiarity with monitoring, alerting, incident response, and root cause analysis • Proficiency in at least one scripting language: Python, Bash, or Go • Understanding of Kubernetes, Docker, and microservices concepts • Knowledge of incident response and operational best practices • Hands-on experience with MySQL performance tuning, replication, and disaster recovery • Proficiency in SQL and NoSQL database management • Experience with a SaaS, service provider, or distributed systems environment • Familiarity with ITIL/OSS practices and SLO/SLAs • Experience with AWS, GCP, or Azure • Ability to work independently, take ownership, and drive projects from problem discovery through resolution

🏖️ Benefits

• No specific benefits, perks, or compensation extras are stated in the posting • Remote work arrangement in Argentina, Colombia, Costa Rica, or Mexico

Apply Now

Similar Jobs

🕒 August 13

Azumo

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

DevSecOps Engineer securing Azumo’s AI-powered financial infrastructure across Latin America. Hardening cloud systems, managing vulnerabilities, and protecting AI agents while supporting compliance and customer security diligence.

AWS

Cloud

Distributed Systems

🕒 July 29

Teladoc Health

5001 - 10000

🏥 Healthcare

👥 B2C

☁️ SaaS

Site Reliability Engineer at Teladoc Health specializing in Azure observability and monitoring. Ensuring reliability and performance of hybrid cloud infrastructure and services.

🇦🇷 Argentina – Remote

💰 $80M Post-IPO Debt - Teladoc Health on 2016-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Azure

Grafana

Python

Terraform

🕒 July 28

Domus Global

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

DevOps Engineer designing and optimizing AWS infrastructure and CI/CD processes at Nublit. Responsible for automation and continuous improvement in development teams.

🗣️🇪🇸 Spanish Required

AWS

Cloud

Grafana

Jenkins

Kubernetes

Prometheus

Terraform

🕒 July 24

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

DevOps Engineer designing Infrastructure as Code and automating deployment processes for a multicultural engineering team at Software Mind. Collaborating on CI/CD pipelines while ensuring system observability and security compliance.

AWS

Azure

Cloud

Docker

Grafana

Kubernetes

Prometheus

Python

Terraform

🕒 July 3

Cashea

501 - 1000

💳 Fintech

🛍️ eCommerce

👥 B2C

DevSecOps Engineer responsible for application security integration in SDLC for a fintech company. Designing and implementing secure pipelines and collaborating with engineering teams.

🗣️🇪🇸 Spanish Required

SDLC