Senior Database Reliability Engineer (DBRE)

Trumid

United States

Remote

USD 225,000 - 265,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Medical, dental, vision coverage
Collaborative culture
Dynamic on-site office

Job summary

Trumid is seeking a Senior PostgreSQL SRE to own the durability, recoverability, and performance of the Postgres/RDS fleet across all environments. You will work with the databases side of our SRE team to drive reliability through failover drills and observability improvements.

Five+ years in production PostgreSQL at scale, strong grasp of replication, WAL, MVCC, and backup/restore, plus IaC and scripting in Terraform, Python, and Bash, are essential.

Qualifications

  • Five or more years running production PostgreSQL at scale, ideally on RDS/Aurora.
  • Experience with replication, WAL, MVCC and vacuum behavior.
  • Proven ability to own data durability, restore, and DR drills.

Responsibilities

  • Own durability, recoverability, and performance of Postgres/RDS fleet across environments.
  • Verify restore-tested coverage and timed failovers with defined RPO/RTO.
  • Own data lifecycle: retention, access control, encryption posture.
  • Hunt engine pathologies: lock contention, WAL throughput, replica lag, bloat.
  • Develop observability and AI-assisted tooling for health checks and migrations.

Skills

PostgreSQL
RDS/Aurora
WAL mechanics
MVCC
Replication
Backup/Restore
SRE
Terraform
Python
Bash
Linux
Observability
Prometheus/Grafana

Tools

Prometheus
Grafana
Terraform

Job description

About us.

Trumid is a dynamic fintech revolutionizing the landscape of fixed income trading. With intelligent, easy-to-use, electronic solutions, we are rapidly growing and seeking exceptional talent to help redefine the boundaries of technology and finance.

Founded in 2014 by a team of fixed income market experts, Trumid has quickly become one of the top three corporate bond e-trading platforms in the U.S. Today, over 1,300 traders from an extensive and expanding client network of 890+ buy-and sell-side institutions transact on Trumid monthly.

With a rich history of innovation and a unique ability to innovate at scale, we collaborate closely with our clients, iterating quickly toward optimal solutions. With market share and client engagement at all-time highs and our pace of product development faster than ever, this is an exciting and transformative time at Trumid.

Our business model thrives on participation, and so does our company culture. We rely on every team member's contribution to help us accomplish our goals. To succeed at Trumid, you must be curious, passionate about your craft, ambitious, collaborative, and driven. Learn more at www.trumid.com

The opportunity.

This role owns data resilience and continuity, and the scope test is simple: if losing it loses data, or makes data unavailable, its yours. The job exists so that data-layer failure modes - storage contention, replica lag, region loss - are found and retired in drills, not discovered in production. Our reliability doctrine is to assume failure and concentrate statefulness into asmall core of systems proven against specific failure modes. Postgres is the heart of that core: everything around it gets to be disruptible because the data layer is not.

You'd join the databases side of our SRE team, working alongside deep Postgres expertise. Some of what you'd walk into:

  • A production RDS fleet backing a live trading venue, with performance work that goes deep: we've characterized WAL-write contention under concurrent commits down to the fsync level, and are weighing group-commit tuning, dedicated log volumes, and storage-class changes against actual measurements.
  • Disaster recovery as an engineering discipline: automated cross-region failover with promotion measured in minutes, and restore paths (snapshot, point-in-time, logical) validated by timed, documented drills on a fixed cadence.
  • Database observability and AI-assisted tooling: engine performance telemetry exported into Prometheus and Grafana, modular database health-check skills, and an automated reviewer for schema-migration PRs.
What you'll do?
  • Own the durability, recoverability, and performance of the Postgres/RDS fleet across every environment: replication, failover, backup and restore, and storage behavior under load.
  • Make recovery provable: restore-tested coverage of every production database, timed failover drills, and measured RPO/RTO per tier - evidence, not assertion.
  • Own the data lifecycle end to end: retention and cleanup policies that preserve recoverability, access control at the data layer, encryption posture, and knowing where sensitive datalives.
  • Hunt performance pathologies at the engine level: lock contention, WAL throughput, replica lag, bloat, index hygiene, write amplification.
  • Build database observability with the team, and extend our AI-assisted operations tooling (health-check skills, migration PR review).
About you.
  • Five or more years running production PostgreSQL at meaningful scale - ideally on RDS or Aurora - with depth in the internals: replication, WAL mechanics, MVCC and vacuum behavior, query planning and performance.
  • You've owned the full lifecycle of data somewhere, not just the query path: retention, backup and restore, access control, and security posture.
  • Disaster-recovery experience you can talk through concretely - failovers you designed, drills you ran, and what they changed.
  • Infrastructure fluency: infrastructure-as-code (Terraform or similar), scripting (Python, bash, SQL), Linux, and cloud storage/IOPS characteristics.
  • An SRE sensibility: SLOs, blameless postmortems, and a preference for rehearsed over improvised.
  • Clear writing - you leave runbooks and decision records behind you.
Nice to have.
  • Warehouse and pipeline experience (BigQuery, AlloyDB, Kafka-based pipelines) - over time we intend to extend the same reliability guarantees beyond Postgres to every store that holds business data.
  • Experience in regulated or fintech environments; exposure to data classification and compliance review.
  • Kubernetes; Prometheus/Grafana/ELK-style observability stacks.
  • Interest in AI-assisted operations tooling
Employee benefits.
  • Highly competitive compensation
  • Fully paid medical, dental and vision coverage
  • Team-oriented and collaborative company culture
  • Lively and dynamic office space with fully stocked kitchen

In compliance with New York City Pay Transparency Law, the base salary range for this role in New York City is between $225,000 - $265,000. This range does not include discretionary bonus or other forms of compensation or benefits offered in connection with this job. Several factors are considered when determining a candidate’s compensation.

Trumid is an equal opportunity employer.

Please note: All communication from Trumid's recruiting team comes from @trumid.com email addresses. We conduct remote interviews via Zoom only. We will never ask you to purchase equipment, download software (other than Zoom), or share sensitive personal information during the hiring process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Network Engineer
Senior Network Engineer

Trumid • United States

On-site
USD 225,000 - 250,000
Competitive compensation
Comprehensive medical/dental/vision
Dynamic office with amenities
Senior Postgres Reliability Engineer (RDS/Aurora)
Senior Postgres Reliability Engineer (RDS/Aurora)

Trumid • United States

Remote
USD 225,000 - 265,000
Competitive compensation
Medical, dental, vision coverage
Collaborative culture
+1
Remote | Senior Database Reliability Engineer — $220,000–$260,000/year
Remote | Senior Database Reliability Engineer — $220,000–$260,000/year

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 220,000 - 260,000
Fully remote
Database Reliability Engineer III
Database Reliability Engineer III

Rent the Runway • New York (NY)

On-site
USD 116,000 - 145,000
Paid Time Off including vacation and family sick leave
Universal Paid Parental Leave
Comprehensive health, vision, dental
+2
Senior Software Engineer, Data Infrastructure (RDBMS)
Senior Software Engineer, Data Infrastructure (RDBMS)

Crypto Pro Network • United States

Remote
USD 200,000 - 220,000
Competitive salary
Equity participation
Remote-first environment
Senior Database Engineer - Postgres
Senior Database Engineer - Postgres

DRW Holdings, LLC. • Chicago (IL), Northern (KY)

Hybrid
USD 150,000 - 225,000
Health insurance
401k with match
Discretionary bonus
+1
Senior Database Reliability Engineer (PostgreSQL, Terraform, AWS)
Senior Database Reliability Engineer (PostgreSQL, Terraform, AWS)

Capital.com • Georgia

On-site
USD 140,000 - 200,000
Competitive salary
Work-life harmony
Generous leave
+1
Senior Data Engineer, Data Platform
Senior Data Engineer, Data Platform

Crypto Pro Network • United States

Remote
USD 190,000 - 220,000
Competitive salary
Remote work flexibility
Participation in equity plan
Senior Software Engineer, Data Platform
Senior Software Engineer, Data Platform

Crypto Pro Network • United States

Remote
USD 190,000 - 220,000
Senior Database Reliability Engineer
Senior Database Reliability Engineer

Artha Nexgen • Northern (KY)

On-site
USD 120,000 - 150,000