Senior Database Reliability Engineer

Jobgether

United States

Hybrid

USD 140,000 - 200,000

Full time

8 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote/Hybrid
Equity (possible)
Flexible PTO
Lunch program
Disability Insurance
Life Insurance
Health benefits
401(k)
Gym/fitness credits
HQ recovery facilities

Job summary

Jobgether is seeking a Senior Database Reliability Engineer in the United States to own the reliability, performance, and cost efficiency of production databases powering a modern healthcare technology platform. You will manage end-to-end database operations across EHR, data, and AI workloads with a focus on preemptive resilience.

You will work with AWS Aurora MySQL, Datadog, Terraform, and HIPAA compliance, guiding platform architecture, implementing runbooks, and optimizing read-topologies and

Qualifications

  • 6+ years in database reliability engineering or similar.
  • Deep MySQL/Aurora MySQL expertise and prod-scale experience.
  • Experience with multi-reader Aurora/RDS clusters at TB scale.
  • Strong observability with Datadog, PMM, Prometheus, or Grafana.
  • Scripting (Python, Bash) and Terraform for automation.
  • AWS RDS/Aurora proficiency and cost-optimization knowledge.
  • On-call rotations, incident response, and runbooks.
  • HIPAA-regulated environment experience is advantageous.
  • Familiarity with security, backup, failover, and DR practices.
  • Excellent cross-team communication and collaboration.

Responsibilities

  • Own reliability, performance, and health of production databases on AWS Aurora MySQL.
  • Manage observability pipelines to Datadog and on-call alerts.
  • Automate safeguards to terminate long-running queries.
  • Analyze execution plans and optimize queries to reduce latency.
  • Review queries before production to gate performance.
  • Educate teams on database hygiene and performance.
  • Create and maintain comprehensive runbooks for incident response.
  • Architect Aurora reader topologies and read-routing strategies.
  • Optimize database cost with right-sizing and Savings Plans.
  • Manage backups, restore testing, failover, and DR capabilities.
  • Standardize schema changes and migrations across teams.
  • Collaborate with platform architecture for new workloads.
  • Handle HIPAA-compliant operations and protect PHI.

Skills

MySQL expertise
Aurora MySQL
AWS RDS/Aurora
Datadog
PMM
Prometheus
Grafana
Python
Terraform
Cloud cost optimization
Incident response

Tools

Datadog
PMM
Terraform

Job description

Senior Database Reliability Engineer based in the United States.

This is a high-impact engineering role responsible for the reliability, performance, availability, and cost efficiency of production databases powering a modern healthcare technology platform.

You will take end-to-end ownership of database operations across EHR, data, and AI workloads, with a strong focus on preventing issues rather than simply reacting to them.

The role combines deep database engineering with observability, automation, incident response, architecture, and cloud cost optimization.

You will partner closely with engineering teams to improve query quality, establish safe migration practices, and ensure scalable database architectures.

You will also play a key role in protecting sensitive healthcare information within a HIPAA-compliant environment.

This is an opportunity to influence platform architecture, engineering standards, and operational practices as the organization continues to scale.

The ideal candidate is a database expert who enjoys solving complex production problems and building systems that remain reliable under growing workloads.

Accountabilities
  • Own the reliability, performance, availability, and operational health of production databases running on AWS Aurora MySQL across EHR, Data, and AI workloads.
  • Manage database observability end to end, maintaining the metrics and alerting pipeline into Datadog and integrating after-hours database alerts into the DevOps on-call rotation.
  • Establish automated safeguards to identify and terminate long-running or runaway queries and provide immediate visibility into database activity across instances.
  • Investigate database performance issues by analyzing execution plans, optimizing queries, re-indexing where appropriate, and reducing unnecessary database load and latency.
  • Review application and EHR queries before production release, serving as a performance gate to prevent inefficient workloads from reaching production.
  • Educate engineering teams on database hygiene, query optimization, and practices that improve reliability and performance.
  • Create and maintain comprehensive database runbooks so first responders can resolve incidents quickly and consistently.
  • Architect and continuously optimize Aurora reader topologies and read-routing strategies based on actual workload requirements.
  • Own database cost efficiency through instance right-sizing, reserved-capacity and Savings Plan strategies, and appropriate storage tiering.
  • Manage replication health, backups, restore testing, failover procedures, and disaster recovery capabilities.
  • Establish safe, repeatable standards for database schema changes and migrations across engineering teams.
  • Partner with platform architecture stakeholders to evaluate and determine the best technical approach for new and existing database workloads.
  • Handle protected health information responsibly while maintaining database operations within a HIPAA-compliant environment.
Requirements
  • 6+ years of experience in database reliability engineering, database administration, database engineering, or a closely related discipline, including ownership of production systems at scale.
  • Deep expertise in MySQL, including query optimization, execution-plan analysis, indexing strategies, and replication; hands‑on Aurora MySQL experience is strongly preferred.
  • Experience operating large, multi‑reader Aurora or RDS clusters at terabyte scale, including read‑routing and connection‑management strategies.
  • Strong knowledge of database observability and monitoring technologies such as Datadog, Percona Monitoring and Management (PMM), Performance Insights, Prometheus, or Grafana.
  • Proficiency with Python, Bash, or a comparable scripting language for automation, along with practical experience using infrastructure‑as‑code tools such as Terraform.
  • Strong AWS operational knowledge, particularly RDS/Aurora, database sizing, reserved capacity, storage options, and cloud cost optimization.
  • Experience with production on‑call rotations, incident response, troubleshooting, and creation of operational runbooks.
  • Strong understanding of database reliability, security, backup, recovery, failover, and disaster‑recovery practices.
  • Excellent communication and collaboration skills, with the ability to educate and influence engineers across multiple teams.
  • Experience with cloud data warehouses such as Snowflake, Databricks, or Redshift is a plus, particularly where analytical workloads can be moved away from OLTP databases.
  • Experience working with HIPAA-regulated or other compliance‑heavy environments involving sensitive data is advantageous.
  • Familiarity with automated query‑remediation approaches such as Percona Toolkit, pt‑kill, statement timeouts, or custom query termination systems is beneficial.
  • Application‑side experience, particularly with PHP, is a plus for effective collaboration on query and application performance.
  • Knowledge of database and cloud security best practices, agile methodologies, and fast‑paced engineering environments is valued.
Benefits
  • Competitive salary and compensation package.
  • Remote/hybrid working environment.
  • Potential equity compensation based on outstanding performance.
  • Flexible PTO.
  • Company-sponsored lunches.
  • Company-paid disability and life insurance.
  • Company-paid family and medical leave.
  • Medical, dental, and vision insurance.
  • Discounted pet insurance.
  • FSA/DCA and commuter benefits.
  • 401(k) retirement plan.
  • Credits toward online fitness classes and gym memberships.
  • Access to an HQ recovery suite featuring a cold plunge, sauna, and shower.
  • Opportunity to work on meaningful healthcare technology and solve complex infrastructure challenges.
  • Collaborative environment focused on smart, sustainable work rather than simply working longer hours.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Database Reliability Engineer
Senior Database Reliability Engineer

Launch Tennessee • United States

Hybrid
USD 150,000 - 210,000
Competitive salaries
Remote/Hybrid environment
Potential equity compensation
+5
Senior Database Reliability Engineer
Senior Database Reliability Engineer

Prompt • United States

Hybrid
USD 180,000 - 240,000
Competitive salary
Remote/hybrid environment
Equity compensation
+3
Senior Database Reliability Engineer
Senior Database Reliability Engineer

Prompt Health • United States

Hybrid
USD 220,000 - 240,000
Remote/hybrid environment
Competitive salaries
401k
Senior Database Reliability Engineer (Remote, Aurora)
Senior Database Reliability Engineer (Remote, Aurora)

Prompt Health • United States

Hybrid
USD 220,000 - 240,000
Remote/hybrid environment
Competitive salaries
401k
Remote Senior DB Reliability Engineer (Aurora MySQL)
Remote Senior DB Reliability Engineer (Aurora MySQL)

Jobgether • United States

Hybrid
USD 140,000 - 200,000
Remote/Hybrid
Equity (possible)
Flexible PTO
+7
Database Reliability Engineer
Database Reliability Engineer

Alex Staff • Georgia

On-site
USD 130,000 - 190,000
Remote work worldwide
Paid vacation (24 days/year)
National holidays (10 days)
+5
Senior Database Engineer, Aurora and Open Source
Senior Database Engineer, Aurora and Open Source

Amazon Web Services (AWS) • Redmond (WA)

On-site
USD 146,000 - 197,000
Health insurance
401(k) matching
Paid time off
Senior Database Engineer, Aurora and Open Source
Senior Database Engineer, Aurora and Open Source

Amazon Web Services (AWS) • East Palo Alto (CA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
Sr Database Engineer
Sr Database Engineer

Cox • Draper (UT)

On-site
USD 102,000 - 169,000
Sr Database Engineer
Sr Database Engineer

Cox Automotive Inc. • Draper (UT)

On-site
USD 102,000 - 169,000