Senior Site Reliability Engineer, Databases

🔥 20 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Vultr

Vultr

201 - 500 employees

Founded 2014

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

💰 $329M Debt Financing - Vultr on 2025-06

Artificial Intelligence • B2B • Hardware

Vultr is a global cloud infrastructure provider offering on-demand virtual machines, bare-metal servers, GPU-accelerated instances, managed databases, object and block storage, Kubernetes, and networking services. The platform emphasizes AI and HPC workloads with a broad selection of AMD and NVIDIA GPUs, fast networking, and 32+ data center regions, plus a marketplace of deployable apps and developer-friendly APIs. Vultr targets developers and businesses seeking affordable, scalable, and compliant cloud compute and storage alternatives to hyperscalers.

📋 Description

• Own and evolve comprehensive monitoring and alerting for MySQL InnoDB Clusters and PostgreSQL infrastructure across multiple datacenters • Establish and execute a quarterly backup verification and disaster recovery testing program across all database systems • Serve as first-tier on-call responder for database incidents, executing documented runbooks and escalating to senior engineering when architecture-level decisions are required • Monitor and maintain replication health across databases • Own database user lifecycle management, including provisioning, deprovisioning, access audits, and role-based access control • Execute and track security compliance remediation, including pen-test findings, GRC audit requirements, and encryption-at-rest verification • Manage operational health of the Debezium/Kafka Connect data pipeline in coordination with the Kafka infrastructure team • Build and maintain Puppet profiles for database infrastructure configuration management • Write PHP, Python, or Go automation tooling to reduce operational toil • Develop and maintain runbooks, operational documentation, and disaster recovery procedures • Partner with the Senior Platform Engineer (Databases) as a peer, reviewing work, sharing on-call, and splitting ownership of database reliability across production systems

🎯 Requirements

• 7+ years of experience in Site Reliability Engineering, DevOps, or Database Operations roles in production environments at scale • Deep operational expertise with MySQL in production — InnoDB Cluster, Group Replication, MySQL Router, and ProxySQL — with strong troubleshooting skills for replication, performance, and reliability issues • Production experience with PostgreSQL — replication, high availability, performance tuning, and operational management • Strong proficiency with configuration management tools (Puppet preferred) and infrastructure-as-code practices • Experience with database backup tools (xtrabackup, pg_dump, mysqldump) and disaster recovery procedures • Proficiency in PHP and Python or Go for automation, tooling, and integration with existing codebases • Experience with database observability — Prometheus exporters, Grafana dashboards, alerting frameworks, and SLO/error-budget methodology • Familiarity with Kafka Connect, Debezium, or similar change-data-capture pipelines • Strong incident response skills with experience in on-call rotations, including post-incident review and remediation • Excellent communication skills and ability to collaborate across engineering teams as a senior peer • Must be residing in one of the listed U.S. states • Must be legally authorized to work in the United States

🏖️ Benefits

• 100% company-paid insurance premiums for employee medical, dental and vision plans • 401(k) plan that matches 100% up to 4%, with immediate vesting • Professional Development Reimbursement of $2,500 each year • 11 Holidays + Paid Time Off Accrual + Rollover Plan • Increased PTO at 3 year and 10 year anniversary • 1 month paid sabbatical every 5 years • Anniversary Bonus each year • $500 stipend for remote office setup in first year + $400 each following year • Internet reimbursement up to $75 per month • Gym membership reimbursement up to $50 per month • Company paid Wellable subscription

Apply Now

Similar Jobs

🔥 23 hours ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Site Reliability Engineer supporting Cisco-owned Splunk’s FedRAMP cloud platform. Testing features, automating infrastructure, and leading customer incidents on overnight remote shifts.

🕒 Yesterday

CACI International Inc

10,000+ employees

🎖️ Defense

🏛️ Government

🔒 Cybersecurity

Operating System Deployment Engineer managing secure Windows and Windows Server images for CACI’s DoD enterprise IT services. Automating deployments and supporting physical and virtual infrastructure across 187 bases.

🇺🇸 United States – Remote

💵 $75.2k - $158.1k / year

🔥 Funding within the last year

💰 $500M Post-IPO Debt on 2026-02

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 Yesterday

PingWind Inc. (SDVOSB)

51 - 200

💼 Consulting

📦 Logistics

🏥 Healthcare

DevSecOps Engineer securing cloud-native software delivery for PingWind, a federal government services provider. Building CI/CD security, automating controls, and supporting vulnerability remediation.

🕒 Yesterday

NCSA College Recruiting

1001 - 5000

⚽ Sports

🎯 Recruiter

📚 Education

DevOps/SRE improving AWS infrastructure, CI/CD, and application reliability. Supporting technology tools for student athletes, coaches, and event operators.

🕒 Yesterday

IMG Academy

501 - 1000

📚 Education

⚽ Sports

🏨 Hospitality

DevOps/SRE engineer optimizing CI/CD, AWS infrastructure, and observability. Supporting technology tools for student athletes, coaches, and event operators.