Senior Site Reliability Engineer, Databases

Vultr

United States

Remote

USD 125,000 - 135,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Insurance premiums
401(k) matching
Professional development reimbursement
Paid time off + Sabbatical program
Remote office stipend
Internet reimbursement
Gym membership reimbursement
Wellable subscription

Job summary

Vultr is seeking a Senior Site Reliability Engineer, Databases to strengthen the Platform Engineering team. You will own database monitoring, backups, DR testing, on-call response, replication health, and access management across MySQL InnoDB Clusters and PostgreSQL.

You will implement automation in PHP, Python or Go, and work with Debezium/Kafka Connect alongside the Kafka team to ensure reliable database infrastructure powering a global cloud platform.

Qualifications

  • 7+ years in SRE/DevOps or DB Operations in production at scale.
  • Deep ops with MySQL InnoDB Cluster, Group Replication, MySQL Router, ProxySQL.
  • PostgreSQL replication, HA, performance tuning, and ops management.
  • Experience with Puppet and IaC.
  • Experience with database backup tools and DR procedures.
  • Proficiency in PHP and Python or Go for automation.
  • Familiarity with Kafka Connect, Debezium, or similar pipelines.
  • Strong incident response and on-call rotation experience.
  • Excellent cross-team communication.

Responsibilities

  • Own monitoring and alerting for MySQL InnoDB Clusters and PostgreSQL across datacenters.
  • Run quarterly backup verifications and DR testing.
  • Serve as first-tier on-call for database incidents and escalate as needed.
  • Monitor replication health and manage DB user lifecycles and RBAC.
  • Ensure encryption-at-rest and security remediation with audits.
  • Maintain Debezium/Kafka Connect data pipeline health.
  • Develop automation tooling to reduce toil using PHP/Python/Go.
  • Document runbooks and disaster recovery procedures.
  • Collaborate with Senior Platform Engineer on database reliability.

Skills

SRE/DevOps experience
MySQL InnoDB Clusters
PostgreSQL administration
Puppet/Infrastructure-as-code
Backup & DR planning
Automation (PHP/Python/Go)
Kafka Connect/Debezium
On-call incident response
Cross-team collaboration

Tools

Xtrabackup
pg_dump
mysqldump
Prometheus/Grafana

Job description

Who We Are

Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation. Founded by David Aninowsky and self-funded for over a decade, Vultr has grown to become the world’s largest privately-held cloud infrastructure company.

Vultr Cares
  • 100% company-paid insurance premiums for employee medical, dental and vision plans.

  • 401(k) plan that matches 100% up to 4%, with immediate vesting

  • Professional Development Reimbursement of $2,500 each year

  • 11 Holidays + Paid Time Off Accrual + Rollover Plan

  • Commitment matters to Vultr! Increased PTO at 3 year and 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year

  • $500 stipend for remote office setup in first year + $400 each following year

  • Internet reimbursement up to $75 per month

  • Gym membership reimbursement up to $50 per month

  • Company paid Wellable subscription

Join Vultr

Vultr is seeking a highly skilled Senior Site Reliability Engineer, Databases to join our Platform Engineering team. You will be the reliability and operational backbone for Vultr's database infrastructure spanning MySQL InnoDB Clusters, PostgreSQL, and other database technologies. Working alongside our Senior Platform Engineer (Databases), you will own database monitoring, backup verification, disaster recovery testing, on-call incident response, replication health, access management, and compliance remediation. This is a peer-level role where you will apply SRE methodology - error budgets, runbooks, automation-first thinking, and toil reduction - specifically to database systems that power a global cloud platform serving millions of customers. You bring deep operational database expertise to complement our existing schema and DDL strengths, and you are comfortable writing PHP, Python, or Go to automate and instrument everything you build.

Key Responsibilities

  • Own and evolve comprehensive monitoring and alerting for MySQL InnoDB Clusters, and PostgreSQL infrastructure across multiple datacenters

  • Establish and execute a quarterly backup verification and disaster recovery testing program across all database systems

  • Serve as first-tier on-call responder for database incidents, executing documented runbooks and escalating to senior engineering when architecture-level decisions are required

  • Monitor and maintain replication health across databases

  • Own database user lifecycle management - provisioning, deprovisioning, access audits, and role-based access control across all database systems

  • Execute and track security compliance remediation including pen-test findings, GRC audit requirements, and encryption-at-rest verification

  • Manage operational health of the Debezium/Kafka Connect data pipeline in coordination with the Kafka infrastructure team

  • Build and maintain Puppet profiles for database infrastructure configuration management and write PHP, Python, or Go automation tooling to reduce operational toil

  • Develop and maintain runbooks, operational documentation, and disaster recovery procedures for all database systems

  • Partner with the Senior Platform Engineer (Databases) as a peer - reviewing each other's work, sharing on-call, and splitting ownership of database reliability across production systems

Qualifications

  • 7+ years of experience in Site Reliability Engineering, DevOps, or Database Operations roles in production environments at scale

  • Deep operational expertise with MySQL in production - InnoDB Cluster, Group Replication, MySQL Router, and ProxySQL - with strong troubleshooting skills for replication, performance, and reliability issues

  • Production experience with PostgreSQL - replication, high availability, performance tuning, and operational management

  • Strong proficiency with configuration management tools (Puppet preferred) and infrastructure-as-code practices

  • Experience with database backup tools (xtrabackup, pg_dump, mysqldump) and disaster recovery procedures

  • Proficiency in PHP and Python or Go for automation, tooling, and integration with existing codebases

  • Experience with database observability - Prometheus exporters, Grafana dashboards, alerting frameworks, and SLO/error-budget methodology

  • Familiarity with Kafka Connect, Debezium, or similar change-data-capture pipelines

  • Strong incident response skills with experience in on-call rotations, including post-incident review and remediation

  • Excellent communication skills and ability to collaborate across engineering teams as a senior peer

Compensation

$125,000 - $135,000

We are currently accepting applications from candidates residing in the following states: Alabama, Arizona, Colorado, Connecticut, Florida, Georgia, Idaho, Illinois, Indiana, Iowa, Kentucky, Louisiana, Maryland, Massachusetts, Michigan, Minnesota, Missouri, Montana, Nebraska, Nevada, New Jersey, New Mexico, New York, North Carolina, Ohio, Oklahoma, Pennsylvania, Rhode Island, South Carolina, Tennessee, Texas, Utah, Vermont, Virginia, Wisconsin.

Inclusion & Privacy

We are an equal opportunity employer and are committed to creating an inclusive environment for all employees. We welcome applications from individuals of all backgrounds and experiences, and we prohibit discrimination based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected status under applicable laws. Vultr will consider qualified applicants with arrest or conviction records in accordance with applicable laws and will not conduct a background check until after an offer of employment has been extended and accepted.

We also take your privacy seriously. We handle personal information responsibly and follow applicable laws, including U.S. privacy rules and India's Digital Personal Data Protection Act, 2023. Your data is used only for legitimate business purposes and is protected with proper security measures.

Where allowed by law, applicants may request details about the data we collect, access or delete their information, withdraw consent for its use, and opt out of nonessential communications. For more details, please see our Privacy Policy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer, Databases
Senior Site Reliability Engineer, Databases

WebHosting • Northern (KY)

Hybrid
USD 125,000 - 135,000
Premium health insurance
401(k) match up to 4%
Professional development reimbursement
+3
Manager, Production Site Reliability Engineering
Manager, Production Site Reliability Engineering

Vultr • United States

Remote
USD 140,000 - 160,000
Company-paid insurance
401(k) match
Professional development reimbursement
+4
Manager, Production Site Reliability Engineering
Manager, Production Site Reliability Engineering

WebHosting • Northern (KY)

Hybrid
USD 140,000 - 160,000
401(k) match up to 4%
Paid holidays & PTO
Professional development reimbursement
+1
Senior Linux System Administrator
Senior Linux System Administrator

Vultr • United States

Remote
USD 80,000 - 100,000
Senior Technical Support Engineer
Senior Technical Support Engineer

Vultr • Town of Montana (WI)

On-site
USD 90,000 - 110,000
Company-paid health premiums
401(k) with matching
Professional development reimbursement
+6
Technical Support Specialist
Technical Support Specialist

Vultr • United States

On-site
USD 34,440 - 41,328
100% company-paid insurance premiums
401(k) plan with matching
Professional Development Reimbursement
+4
Senior Talent Acquisition Specialist
Senior Talent Acquisition Specialist

Vultr • United States

On-site
USD 80,000 - 95,000
100% company-paid insurance premiums
401(k) with matching
Professional Development Reimbursement
+5
Senior Linux System Administrator
Senior Linux System Administrator

Vultr • West Palm Beach (FL)

Remote
USD 80,000 - 100,000
Remote-first company
Internet reimbursement
Education reimbursement of up to $2,?5
+2
Staff Platform Engineer
Staff Platform Engineer

Webhosting • Northern (KY)

On-site
USD 145,000 - 160,000
Health insurance
401(k) match
Professional development
+6
Senior Platform Engineer
Senior Platform Engineer

WebHosting • Northern (KY)

Hybrid
USD 125,000 - 135,000
100% company-paid health insurance
401(k) matching 100% up to 4%
Professional development reimbursement
+6