Search Remote Jobs

Senior Lead Database Reliability Engineer

🕒 July 17

đŸ‡ș🇾 United States – Remote

đŸ’” $168k - $210k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ‘» Ghost score 7%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of DraftKings Inc.

DraftKings Inc.

1001 - 5000 employees

Founded 2012

🎼 Gaming

⚜ Sports

đŸ‘„ B2C

Gaming ‱ Sports ‱ B2C

DraftKings Inc. is a global company known for providing innovative products and experiences primarily in the sports betting and fantasy sports sectors. The company boasts a strong presence across multiple countries, aiming to deliver exceptional customer moments and overcoming challenges through teamwork and persistence. At DraftKings, innovation in engineering, analytics, and product development is key, with a focus on creating unforgettable customer experiences in sportsbook and casino operations. The company emphasizes a dynamic work culture, inclusion, equity, and global collaboration within its diverse teams.

📋 Description

‱ Drive the technical roadmap for database reliability across PostgreSQL, MySQL, MongoDB, Redis, ScyllaDB, Aerospike, and managed cloud services, while shaping architecture for high availability, replication, partitioning, storage, and connection management ‱ Design and build automation first database platforms by developing Kubernetes operators, infrastructure as code, GitOps workflows, and production quality tooling in Go or Python to automate provisioning, failover, backups, schema migrations, and lifecycle management ‱ Lead operational excellence by defining service level objectives, monitoring database health, capacity, and performance, eliminating recurring reliability issues, validating backup and recovery processes, and leading critical production incidents through resolution and continuous improvement ‱ Optimize database performance and cost across cloud and on premises environments by driving capacity planning, resource efficiency, storage optimization, workload consolidation, and performance tuning for large scale systems ‱ Partner closely with application engineering teams to establish safe database practices, including schema reviews, migration strategies, query optimization, connection management, and zero downtime deployment processes ‱ Leverage AI to improve engineering productivity and database operations through intelligent observability, anomaly detection, root cause analysis, documentation, predictive insights, and evaluation of AI generated code to ensure reliability and security ‱ Mentor engineers across the organization by sharing best practices, leading design and code reviews, influencing technical direction, supporting hiring efforts, and raising the overall maturity of database reliability engineering.

🎯 Requirements

‱ At least 6 years of experience in Database Reliability Engineering, Database Platform Engineering, or Site Reliability Engineering with a strong database focus ‱ Deep expertise in at least one major relational database, preferably PostgreSQL, along with operational experience supporting technologies such as MySQL, MongoDB, Redis, ScyllaDB, Aerospike, Aurora, Cloud SQL, and other managed cloud database services ‱ Strong experience building and operating stateful workloads on Kubernetes using technologies such as StatefulSets, Persistent Volumes, database operators, Terraform, Pulumi, FluxCD, ArgoCD, GKE, and EKS ‱ Hands on software development experience using Go or Python to create automation, platform tooling, Kubernetes controllers, APIs, and infrastructure ‱ A data driven, automation first mindset with proven experience improving reliability through observability, monitoring, service level objectives, capacity planning, performance optimization, and self service engineering solutions ‱ Practical experience using AI tools such as Claude, GitHub Copilot, Cursor, MCP, or similar technologies to improve design, coding, documentation, troubleshooting, and operational workflows while applying sound engineering judgment to validate AI generated outputs ‱ Excellent leadership and communication skills with a track record of mentoring engineers, influencing architectural decisions, collaborating across engineering teams, producing clear technical documentation, and driving continuous improvement in highly available production environments.

đŸ–ïž Benefits

‱ plus bonus ‱ equity ‱ benefits as applicable

Apply Now

Similar Jobs

🕒 July 17

Upstart

1001 - 5000

🚘 Automotive

đŸ’Œ Consulting

đŸ„ Healthcare

Lead the SRE team at Upstart to enhance product reliability and operational maturity. Drive incident management, observability, and reliability engineering improvements across the company.

🕒 July 17

Siemens Healthineers

10,000+ employees

đŸ„ Healthcare

⚕ Healthcare Insurance

🧬 Biotechnology

Network Engineer responsible for cloud solutions operations at Varian. Collaborating across teams to enhance, optimize, and maintain cloud computing capabilities.

🕒 July 17

Siemens Healthineers

10,000+ employees

đŸ„ Healthcare

⚕ Healthcare Insurance

🧬 Biotechnology

Network Engineer managing cloud operations and support for Varian’s cloud solutions. Enhancing, optimizing, and maintaining computing capabilities across the global cloud solution.

🕒 July 16

Miris

11 - 50

☁ SaaS

đŸ„œ AR/VR

đŸ€ B2B

Site Reliability Engineer building scalable platforms for 3D/4D content delivery at Miris. Collaborating with teams to ensure system reliability and performance across AR/VR devices.

🕒 July 16

Epic Kids

11 - 50

📚 Education

đŸ‘„ B2C

đŸ“± Media

Senior Site Reliability Engineer driving reliability and stability for Epic Kids' GCP infrastructure. Collaborating with product and data teams to ensure platform efficiency and security.