Database Reliability Engineer

Job not on LinkedIn

đŸ”„ 3 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Alex Staff Agency

Alex Staff Agency

11 - 50 employees

🎯 Recruiter

đŸ€ B2B

Recruitment ‱ B2B

Alex Staff Agency is a recruitment agency specializing in connecting IT professionals with companies around the globe. The agency offers a variety of services including career counseling for candidates, helping them find job opportunities in countries they love, and assisting employers in hiring skilled talent. With over 20 years of experience in the field, Alex Staff Agency has built a broad network across more than 150 markets, facilitating successful relocations and onboarding for numerous candidates. The agency prides itself on supporting candidates throughout the hiring process and beyond, offering guidance and expertise every step of the way.

📋 Description

‱ Own production PostgreSQL reliability: HA design, Patroni, PgBouncer, replication, failover, upgrades, vacuum/bloat control, query tuning, locks, indexes, capacity, backups, PITR, and restore validation. ‱ Improve disaster recovery and operational evidence: tested restores, documented recovery paths, measurable RTO/RPO targets, runbooks, and safe maintenance plans. ‱ Support the wider database estate: ClickHouse, MongoDB, and Redis. ‱ Troubleshoot incidents, review access and data-safety changes, improve monitoring, and learn the production ClickHouse patterns already in use. ‱ Automate DBA workflows with Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts, and reproducible runbooks for provisioning, grants, backups, restores, health checks, and ownership metadata. ‱ Help build DBaaS-style self-service capabilities so engineering teams can request databases, access, credentials, and operational checks with less manual DBA intervention. ‱ Improve observability and incident response through Grafana, metrics, logs, SLOs, alert rules, Opsgenie routing, and clear communication during production issues.

🎯 Requirements

‱ Deep hands-on PostgreSQL experience in business-critical production environments, typically 5+ years or equivalent depth. ‱ Strong understanding of PostgreSQL internals and operations: MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing. ‱ Proven experience with highly available databases and the ability to reason about quorum, split-brain risk, failover, rollback, and recovery. ‱ Strong Linux and infrastructure fundamentals: systemd, networking, storage, filesystems, CPU/memory/disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting. ‱ Automation skills with Ansible and scripting. Terraform/OpenTofu, GitLab CI/CD, and merge-request based delivery are strong advantages. ‱ Ability to support more than one database engine. You do not need to be a ClickHouse expert on day one, but you must be ready to learn it quickly and take responsibility for it. ‱ Practical use of AI engineering assistants such as Claude and Codex. We expect you to use them to improve speed and quality, while personally verifying generated SQL, commands, scripts, and operational conclusions. ‱ English - upper-intermediate or higher - to ensure clear communication of progress within the teams. ‱ Nice to Have: ‱ ClickHouse operations: replication, Keeper/ZooKeeper, MergeTree engines, distributed DDL, grants, row policies, backups, query troubleshooting, and cluster recovery. ‱ MongoDB replica sets and Percona Backup for MongoDB. ‱ Redis/Sentinel and broker/cache failure modes. ‱ Database observability, SLOs, golden signals, alert tuning, and executable incident runbooks. ‱ Building internal platforms, self-service portals, or DBaaS workflows for engineering teams.

đŸ–ïž Benefits

‱ A focus on professional development. ‱ Interesting and challenging projects. ‱ Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide. ‱ Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves. ‱ Compensation for private medical insurance. ‱ Co-working and gym/sports reimbursement. ‱ Budget for education. ‱ The opportunity to receive a reward for the most innovative idea that the company can patent.

Apply Now

Similar Jobs

đŸ”„ 3 hours ago

Mirantis

501 - 1000

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Senior AI Deployment Engineer at Mirantis, specializing in AI infrastructure deployment using Kubernetes and NVIDIA-certified hardware. Collaborate with international teams and optimize cloud-based AI solutions.

Cloud

Distributed Systems

JavaScript

Kubernetes

Linux

Microservices

Open Source

OpenStack

Python

Go

đŸ”„ 14 hours ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏱 Enterprise

Customer Reliability Engineer engaging with enterprise customers to adapt SRE practices across Isovalent products. Focusing on customer architecture and configurations while providing world-class support and incident resolution.

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Linux

Terraform

🕒 Yesterday

COLIBRIX ONE

51 - 200

💳 Fintech

🏩 Banking

🔌 API

DevOps Engineer at COLIBRIX ONE managing infrastructure for AI-powered payment technologies and collaborating on innovative fintech solutions in a fast-growing team environment.

đŸ—ŁïžđŸ‡·đŸ‡ș Russian Required

AWS

Cloud

DNS

EC2

Kubernetes

Linux

TCP/IP

Terraform

🕒 Yesterday

Nord Security

1001 - 5000

🔒 Cybersecurity

☁ SaaS

đŸ€ B2B

Site Reliability Engineer at NordVPN Apps managing infrastructure and enhancing security. Collaborate with teams to optimize and secure services while ensuring availability and reliability.

đŸ‡”đŸ‡± Poland – Remote

đŸ’” zƂ23.3k - zƂ34k / month

💰 $100M Private Equity Round - Nord Security on 2023-09

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Grafana

Linux

Python

🕒 Yesterday

Nord Security

1001 - 5000

🔒 Cybersecurity

☁ SaaS

đŸ€ B2B

Senior Site Reliability Engineer responsible for ensuring content accessibility across a global network for millions of users. Involves load balancing, traffic shaping and troubleshooting at protocol levels.

đŸ‡”đŸ‡± Poland – Remote

đŸ’” zƂ23.3k - zƂ34k / month

💰 $100M Private Equity Round - Nord Security on 2023-09

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

DNS

Docker

HAProxy

Linux

NGINX

Python

SaltStack