Senior Database Administrator

hardrockdigital

United States

Hybrid

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible vacation allowance
Hybrid / remote working environment
Employer-sponsored training and conference attendance

Job summary

Hard Rock Digital is looking for a Senior Database Administrator in the United States to ensure the reliability of their CockroachDB and PostgreSQL clusters. The role involves diagnosing database issues, improving system performance, and working with advanced tools like Go and Python.

In this role, you will operate multi-region databases, lead incident analyses, and contribute to building a robust data platform. Candidates need extensive experience in database management and distributed systems.

Qualifications

  • 5+ years operating production relational databases as a DBA or similar role.
  • Deep production experience with PostgreSQL and CockroachDB.
  • Understanding of distributed systems, consensus, and transaction models.

Responsibilities

  • Operate and evolve multi-region CockroachDB and PostgreSQL clusters.
  • Diagnose and handle database contention and performance issues.
  • Lead incident analysis and drive blameless postmortems.

Skills

PostgreSQL
CockroachDB
Python
Go
Distributed systems knowledge
UNIX/Linux administration

Tools

Terraform
Prometheus
Grafana

Job description

What are we building?

Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We're building a team that resonates with passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer's interactions, experiences, behavior, and insights, and we strive to ensure we always act authentically.

Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, the new Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We're taking that foundation of success and bringing it to the digital space. Ready to join us?

What's the position?

Every bet and payout at Hard Rock Digital is a database transaction that has to be fast and accurate. A casino game resolves in milliseconds; settling a big sporting event moves millions of balances at once, and none can be wrong.

We're hiring a Senior Database Administrator to own production reliability for our CockroachDB and PostgreSQL fleet. You'll diagnose contention, range hotspots, retry storms, plan regressions, and cross-region latency in systems that move money at scale. You'll read execution plans and reason about the optimizer that produced them, inspect database internals when behavior doesn't match the docs, write Go and Python tooling that makes the next incident faster to solve, and set standards that application engineers rely on under pressure.

Claude Code, Codex, and similar AI harnesses are part of how this team works. You'll use them to investigate incidents, generate probes, summarize evidence, and turn one-off debugging into durable tooling. The goal is to spend less time on toil and more on the work that compounds: better defaults, safer migrations, and tooling that makes the next incident shorter than the last.

What You'll Do
Reliability & Performance
  • Operate and evolve multi-region CockroachDB clusters and PostgreSQL instances across production, staging, and development.
  • Investigate contention, serialization retries (SQLSTATE 40001), and range hotspots in the transaction paths behind bets, wagers, settlements, and payouts.
  • Diagnose leaseholder placement, monotonic-key write pressure, and cross-region latency that turn one query into several network round trips.
  • Trace a slow query to its execution plan, the plan to optimizer behavior, and the behavior to the schema or application call path that caused it.
  • Catch plan regressions after statistics changes, schema changes, data growth, or releases, then turn the fix into a safer rollout pattern.
  • Assess schema-change and backfill risk before it reaches large tables: lock behavior, retry pressure, capacity impact, and rollback path.
  • Plan capacity, scaling events, and version upgrades that customers don't feel.
Observability & Incident Response
  • Build database observability with Prometheus, Grafana, Mimir, Loki, and Snowflake: dashboards backed by real queries, alerting on leading indicators, and SLO tracking.
  • Connect database symptoms to application behavior and business events so alerts point at causes.
  • Join the on-call rota for the data layer, lead incident analysis when the database is part of the failure, and drive blameless postmortems toward durable fixes.
  • Instrument the system to surface degradation before it becomes an incident.
Automation & Tooling
  • Build Go and Python tools that make incidents easier to understand: log collectors, explain-plan analyzers, migration checks, capacity models, and runbook generators.
  • Automate provisioning, configuration, backup, and recovery with Terraform and other infrastructure-as-code tools.
  • Work with platform engineering on CI/CD, deployment safety, and change management for database-touching services.
AI-Augmented Engineering
  • Use Claude Code, Codex, and similar harnesses in daily work to investigate incidents, generate probes, draft runbooks, and write tooling.
  • Keep the harness grounded in logs, traces, metrics, and source code, and verify its output against production facts before you ship it.
  • Evaluate AI-assisted observability and anomaly detection, and share what works with the team.
Security & Compliance
  • Implement and audit access controls, authentication, authorization, and encryption in transit and at rest.
  • Support data-protection, audit-logging, and retention requirements for a regulated gaming platform.
Standards & Mentorship
  • Build a database center of excellence: documentation, standards, and reusable patterns that other teams build against.
  • Mentor engineers on database fundamentals, distributed-systems behavior, and operational discipline.
  • Read CockroachDB and PostgreSQL internals when a problem needs a code-level answer, and bring what you learn back to the team.
What are we looking for?
  • 5+ years operating production relational databases as a DBA, Database Reliability Engineer, or Data Platform Engineer.
  • Deep production experience with PostgreSQL, CockroachDB, or another relational database management system, with fluency in the tradeoffs behind isolation, consensus, locality, and query planning.
  • You understand how a SQL engine works under the hood: MVCC, the Volcano iterator execution model, and cost-based optimizer frameworks such as Cascades. You apply that knowledge to indexing strategy and to execution plan analysis.
  • Working knowledge of distributed-systems fundamentals: consensus (Raft and Paxos), distributed transactions, consistency models, and failure modes.
  • You write tools and automation in Python or Go, and you can debug and suggest fixes to line-of-business code in Java or a similar language.
  • Experience with infrastructure-as-code and a monitoring and observability stack such as Grafana or Datadog.
  • UNIX/Linux administration with shell scripting, comfort on a cloud platform (AWS preferred, Azure, or GCP welcome), and willingness to join an on-call rota.
  • Clear technical writing and speech. You'll work across teams.
What You've Done

The strongest candidates show evidence here, not a checklist.

  • Shipped Go or Python tooling for database operations, performance analysis, migration safety, or incident response that other engineers then used.
  • Investigated production incidents end-to-end using metrics, logs, query fingerprints, execution plans, and source-level reasoning.
  • Owned a database platform through on-call, post-incident repair, standards, and cross-team adoption.
  • Used AI coding agents in real engineering work, with review habits that keep generated output tied to evidence.
  • Run CockroachDB in production.
  • Contributed to an open-source database such as PostgreSQL.
  • Worked with streaming infrastructure such as Kafka and Materialize.
  • Worked with a modern data warehouse or analytics platform such as ClickHouse Cloud or Snowflake.
  • Practiced SRE along the lines of the Google SRE books.
What’s in it for you?

We offer our employees more than just competitive compensation. Our team benefits include:

  • Competitive pay and benefits
  • Flexible vacation allowance
  • A hybrid / remote working environment
  • Startup culture backed by a secure, global brand
  • Opportunity to build products enjoyed by millions as part of a passionate team
  • Team gatherings around the US and Europe
  • Employer-sponsored training and conference attendance
  • Opportunity to work in an AI-first environment with access to the tools you need to excel at your job
Roster of Uniques

We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Database Reliability Engineer
Database Reliability Engineer

Alex Staff • Georgia

On-site
USD 130,000 - 190,000
Remote work worldwide
Paid vacation (24 days/year)
National holidays (10 days)
+5
Senior Engineer - DevOps
Senior Engineer - DevOps

Hard Rock Digital • United States

Hybrid
USD 160,000 - 210,000
Hybrid/Remote work options
Competitive compensation
Staff+ Software Engineer, Databases
Staff+ Software Engineer, Databases

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Manager, Engineering (Production Orchestration)
Manager, Engineering (Production Orchestration)

Cockroach Labs • New York (NY)

On-site
USD 194,000 - 258,000
Stock Options
Medical Insurance
Vision Insurance
+9
Lead Analyst - Insights
Lead Analyst - Insights

hardrockdigital • Nevada (IA)

Hybrid
USD 90,000 - 130,000
Competitive pay
Training and development
Flexible vacation
+3
Lead Analyst - Insights
Lead Analyst - Insights

Hard Rock Digital • Town of Florida (NY)

Hybrid
USD 110,000 - 150,000
Competitive pay
Training & development
Flexible vacation
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
AI Forward Deployed Engineer - CRM & Marketing
AI Forward Deployed Engineer - CRM & Marketing

Hard Rock Digital • United States

Hybrid
USD 100,000 - 130,000
Competitive pay and benefits
Flexible vacation allowance
Hybrid / remote working environment
+1
Member of Technical Staff (Production Services)
Member of Technical Staff (Production Services)

The available sources do not contain information about the company name for rounx.com. • New York (NY)

Hybrid
USD 140,000 - 230,000
Stock Options
Medical Insurance
Vision Insurance
+9
Senior Database Reliability Engineer (PostgreSQL, Terraform, AWS)
Senior Database Reliability Engineer (PostgreSQL, Terraform, AWS)

Capital.com • Georgia

On-site
USD 140,000 - 200,000
Competitive salary
Work-life harmony
Generous leave
+1