Database Reliability Engineer

Alex Staff

Georgia

On-site

USD 130,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work worldwide
Paid vacation (24 days/year)
National holidays (10 days)
Unlimited sick leave
Private medical insurance
Co-working and gym allowance
Education budget
Innovation patent reward

Job summary

Alex Staff seeks a Senior Database Reliability Engineer to own production PostgreSQL reliability and help automate DBA workflows across ClickHouse, MongoDB, and Redis. You will design HA, manage failover, backups, and upgrades while learning the ClickHouse environment and ensuring safe operations.

You will automate with Ansible, Terraform/OpenTofu, and GitLab CI/CD, build self-service capabilities, and improve observability with Grafana and runbooks.

Qualifications

  • 5+ years hands-on PostgreSQL in production.
  • Deep PostgreSQL internals: MVCC, WAL, replication, autovacuum, backups.
  • Experience with HA, failover, recovery and quorum considerations.
  • Strong Linux and infrastructure fundamentals.
  • Automation with Ansible; Terraform/OpenTofu; GitLab CI/CD.

Responsibilities

  • Own production PostgreSQL reliability: HA design, replication, upgrades, backups.
  • Improve disaster recovery with tested restores and runbooks.
  • Support ClickHouse, MongoDB, and Redis; troubleshoot incidents.
  • Automate DBA workflows with Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts.
  • Help build DBaaS self-service for databases, access, and checks.
  • Improve observability with Grafana, metrics, logs, SLOs, and incident response.

Skills

PostgreSQL
Linux
Automation
Ansible
Terraform/OpenTofu
GitLab CI/CD
Shell scripting
AI tools use (Claude/Codex)
Multi-DB support

Tools

PostgreSQL
ClickHouse
MongoDB
Redis

Job description

We are hiring a Senior Database Reliability Engineer to join the Infrastructure DBA cell. This is a hands-on production ownership role, not a narrow ticket-processing DBA position. You will keep critical database services reliable, automate repeated work, support engineering teams, and reduce single-person dependency in our PostgreSQL, ClickHouse, MongoDB, and Redis operations.

PostgreSQL is the main requirement. ClickHouse experience is a strong plus, but it is not a day-one blocker. We need a senior engineer with enough database, Linux, automation, and incident-response depth to learn our ClickHouse environment quickly and operate it safely.

Your Responsibilities
  • Own production PostgreSQL reliability: HA design, Patroni, PgBouncer, replication, failover, upgrades, vacuum/bloat control, query tuning, locks, indexes, capacity, backups, PITR, and restore validation.
  • Improve disaster recovery and operational evidence: tested restores, documented recovery paths, measurable RTO/RPO targets, runbooks, and safe maintenance plans.
  • Support the wider database estate: ClickHouse, MongoDB, and Redis. You will troubleshoot incidents, review access and data-safety changes, improve monitoring, and learn the production ClickHouse patterns already in use.
  • Automate DBA workflows with Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts, and reproducible runbooks for provisioning, grants, backups, restores, health checks, and ownership metadata.
  • Help build DBaaS-style self-service capabilities so engineering teams can request databases, access, credentials, and operational checks with less manual DBA intervention.
  • Improve observability and incident response through Grafana, metrics, logs, SLOs, alert rules, Opsgenie routing, and clear communication during production issues.
What Success Looks Like
  • PostgreSQL clusters have tested backup and restore paths, useful dashboards, clear ownership, and documented failover procedures.
  • Repeated DBA tickets become automation or self-service workflows.
  • ClickHouse operational knowledge is no longer a single-person dependency.
  • Database incidents have owners, runbooks, evidence, and measurable recovery paths.
  • Product and engineering teams get database help faster without sacrificing safety, auditability, or reliability.
Why Us

You will work on real production infrastructure used across products.

You will have a direct impact on reliability, incident response, developer experience, and operational resilience.

You will also work in an AI-assisted engineering culture where automation, documentation, Claude, Codex, and careful human verification are part of the daily operating model.

Requirements
What We Expect From You
  • Deep hands-on PostgreSQL experience in business-critical production environments, typically 5+ years or equivalent depth.
  • Strong understanding of PostgreSQL internals and operations: MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing.
  • Proven experience with highly available databases and the ability to reason about quorum, split-brain risk, failover, rollback, and recovery.
  • Strong Linux and infrastructure fundamentals: systemd, networking, storage, filesystems, CPU/memory/disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting.
  • Automation skills with Ansible and scripting. Terraform/OpenTofu, GitLab CI/CD, and merge-request based delivery are strong advantages.
  • Ability to support more than one database engine. You do not need to be a ClickHouse expert on day one, but you must be ready to learn it quickly and take responsibility for it.
  • Practical use of AI engineering assistants such as Claude and Codex. We expect you to use them to improve speed and quality, while personally verifying generated SQL, commands, scripts, and operational conclusions.
  • English - upper-intermediate or higher - to ensure clear communication of progress within the teams.
Nice to Have
  • ClickHouse operations: replication, Keeper/ZooKeeper, MergeTree engines, distributed DDL, grants, row policies, backups, query troubleshooting, and cluster recovery.
  • MongoDB replica sets and Percona Backup for MongoDB.
  • Redis/Sentinel and broker/cache failure modes.
  • Database observability, SLOs, golden signals, alert tuning, and executable incident runbooks.
  • Building internal platforms, self-service portals, or DBaaS workflows for engineering teams.
Benefits
  • A focus on professional development.
  • Interesting and challenging projects.
  • Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide.
  • Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves.
  • Compensation for private medical insurance.
  • Co-working and gym/sports reimbursement.
  • Budget for education.
  • The opportunity to receive a reward for the most innovative idea that the company can patent.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Database Reliability Engineer & Administrator
Database Reliability Engineer & Administrator

TCN • Saint George (UT)

On-site
USD 100,000 - 130,000
Medical Insurance (HDHP with HSA)
Dental Insurance
Vision Insurance
+6
Senior Database Reliability Engineer
Senior Database Reliability Engineer

ClickUp • Town of Poland (NY)

On-site
USD 110,000 - 140,000
Senior Database Reliability Engineer
Senior Database Reliability Engineer

ClickUp • United States

Hybrid
USD 130,000 - 190,000
Senior PostgreSQL Reliability Engineer - Remote & Automation
Senior PostgreSQL Reliability Engineer - Remote & Automation

Alex Staff • Georgia

On-site
USD 130,000 - 190,000
Remote work worldwide
Paid vacation (24 days/year)
National holidays (10 days)
+5
Senior Database Administrator
Senior Database Administrator

hardrockdigital • United States

Hybrid
USD 100,000 - 130,000
Flexible vacation allowance
Hybrid / remote working environment
Employer-sponsored training and conference attendance
Senior Database Reliability Engineer (PostgreSQL, Terraform, AWS)
Senior Database Reliability Engineer (PostgreSQL, Terraform, AWS)

Capital.com • Georgia

On-site
USD 140,000 - 200,000
Competitive salary
Work-life harmony
Generous leave
+1
Senior DB Reliability Engineer - Remote & Flexible Hours
Senior DB Reliability Engineer - Remote & Flexible Hours

Alex Staff • Town of Poland (NY)

On-site
USD 130,000 - 190,000
Fully remote work
Flexible working hours
Vacation 24 days per year
+5
Senior PostgreSQL Database Administrator
Senior PostgreSQL Database Administrator

Jobgether • United States

On-site
USD 130,000 - 150,000
Medical, dental, and vision insurance
401(k) with company match
Health Savings Account contributions
+2
Principal Software Engineer - Postgres
Principal Software Engineer - Postgres

ClickHouse • United States

Hybrid
USD 140,000 - 200,000
Flexible work environment
Healthcare contributions
Company equity
+3
Senior Database Reliability Engineer (DBRE)
Senior Database Reliability Engineer (DBRE)

Tata Consultancy Services • San Jose (CA)

On-site
USD 90,000 - 130,000