Site Reliability Engineer III, DBA

🔥 12 minutes ago

🇺🇸 United States – Remote

💵 $125k - $150k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Backblaze

Backblaze

201 - 500 employees

Founded 2007

🛍️ eCommerce

🏢 Enterprise

💰 $5M Series A on 2012-07

Cloud Storage • eCommerce • Enterprise

Backblaze is a cloud storage company that provides scalable and secure data backup solutions for both businesses and individuals. Their B2 Cloud Storage service offers S3 compatible object storage, allowing users to easily protect and manage their data with transparent pricing. Backblaze specializes in automatic and unlimited backup services for computer systems, ensuring data protection and recovery options for users, while also supporting integration with applications for enhanced functionality.

📋 Description

• Design, deploy, and own highly available database architecture for Vitess (distributed MySQL) and Cassandra • Establish and document operational procedures, runbooks, and escalation guidance for Level 1 and Level 2 SRE Database Engineers • Optimize database performance through query tuning, indexing strategies, schema design, and capacity planning • Own database backup, recovery, replication, and disaster recovery strategies • Perform and validate disaster recovery testing and database recovery procedures • Drive database security, access control, patching, hardening, and compliance practices • Partner with DBA and Data Infrastructure teams on resharding, capacity planning, replication, and architecture decisions for sharded MySQL environments • Support availability and durability of critical services across production environments • Monitor service health using SLIs, SLOs, error budgets, monitoring, logging, and alerting platforms • Participate in on-call rotations, incident response, root cause analysis, and post-incident reviews • Serve as an escalation point for complex database production incidents • Develop automation for operational and database administration tasks • Contribute to monitoring, logging, and alerting frameworks including Prometheus, Grafana, Catchpoint, and ELK • Integrate operational runbooks and incident response workflows with FireHydrant • Work with CI/CD pipelines, configuration management, and infrastructure-as-code tools including Terraform, Ansible, and Jenkins • Develop scripts using Bash, Python, Go, or similar technologies • Operate and troubleshoot containerized production environments using Kubernetes and Docker • Lead Production Readiness Reviews and support operational readiness of new database-backed services • Build training plans, onboarding materials, and technical documentation for Level 1 and Level 2 SRE Database Engineers • Partner with Engineering, Product, Operations, and DBA/Data Infrastructure teams on reliability initiatives • Assist with capacity planning, disaster recovery exercises, database migrations, and infrastructure projects • Work with vendors and service providers to troubleshoot service issues and track SLA performance • Respond to and resolve production database, infrastructure, and service incidents • Troubleshoot and escalate database, Linux, networking, application, and infrastructure issues • Identify recurring issues and develop long-term corrective actions to improve reliability

🎯 Requirements

• 6–8 years of experience in site reliability engineering, systems engineering, infrastructure operations, database engineering, or similar roles, with meaningful experience supporting production database systems • Deep hands-on experience with MySQL and distributed or sharded database systems • Experience with Vitess in a production environment strongly preferred • Experience administering and supporting NoSQL databases such as Cassandra • Experience designing high-availability database architecture, replication topology, backup strategies, and disaster recovery processes • Strong SQL skills, including query performance analysis, indexing, schema design, and troubleshooting • Solid Linux systems administration and troubleshooting skills • Experience with security-focused operations including patching, system hardening, access controls, and vulnerability remediation • Strong understanding of service reliability concepts including monitoring, alerting, incident response, root cause analysis, SLIs, SLOs, and error budgets • Experience working with containers and orchestration platforms including Kubernetes and Docker • Comfortable operating in Kubernetes and Vitess environments using tools such as kubectl, mysqlsh, and Vitess keyspaces • Experience with infrastructure and configuration management technologies including Terraform, Ansible, Jenkins, and HashiCorp products such as Vault and Nomad • Proficiency in at least one scripting language such as Python, Bash, or Go • Experience establishing operational procedures, runbooks, documentation, and escalation processes • Experience mentoring, training, or helping onboard engineers into complex technical environments • Experience in SaaS, cloud services, service provider, or large-scale distributed systems environments preferred • Experience with AWS, GCP, Azure, or similar cloud platforms preferred • Familiarity with ITIL/OSS practices and SLA/SLO management preferred • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience

🏖️ Benefits

• Healthcare for family, including dental and vision • Competitive compensation and 401K • RSU grants for full-time employees • ESPP program • Flexible vacation policy • Maternity & paternity leave • MacBook Pro to use for work, plus a generous stipend to personalize your workstation • Childcare bonus (human children only) • Fertility treatment and support • Learning & development program • Commuter benefits • Culture that supports a healthy work-life balance

Apply Now

Similar Jobs

🔥 24 minutes ago

Shippo

201 - 500

📦 Logistics

📣 Marketing

💼 Consulting

SRE Manager leading Kubernetes, cloud infrastructure, and reliability platforms for Shippo’s global shipping technology. Enabling product teams to deploy and operate scalable services.

🔥 3 hours ago

Climavision

11 - 50

💼 Consulting

📦 Logistics

🤖 Artificial Intelligence

Senior SRE operating Kubernetes, observability, and automated recovery for Climavision’s weather intelligence platform. Driving high availability across Azure, colocation, and edge infrastructure.

🔥 4 hours ago

Synapticure Inc.

11 - 50

🏥 Healthcare

📡 Telecommunications

⚕️ Healthcare Insurance

Senior DevOps and Security Engineer securing AWS, Kubernetes, and software delivery environments. Supporting Synapticure’s virtual neurodegenerative disease care and life sciences research platform.

🔥 4 hours ago

Wursta

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Digital Workplace Deployment Engineer leading Google Workspace migrations and cloud deployment services. Delivering digital transformation, managed services, cybersecurity, and AI solutions at Wursta.

🔥 5 hours ago

Humana

10,000+ employees

🏥 Healthcare

🛡️ Insurance

⚕️ Healthcare Insurance

Senior DevOps Engineer advancing AI/ML DevOps maturity, cloud automation, and CI/CD pipelines at Humana. Improving security, releases, infrastructure provisioning, and software quality.