Site Reliability Engineer III (DBA)

WebHosting

Northern (KY)

Hybrid

USD 125,000 - 150,000

Full time

33 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

ESPP program

Job summary

Backblaze is seeking a Site Reliability Engineer III (DBA) to ensure the stability and reliability of production database systems, primarily Vitess and Cassandra. You will design architecture, develop runbooks, and own on-call incident response with a focus on scalable, observable infrastructure.

You’ll automate routine DB tasks, participate in disaster recovery planning, and collaborate with Kubernetes, CI/CD, and data infrastructure teams.

Qualifications

  • Production database systems (Vitess, Cassandra) experience required.
  • Strong Linux, automation, distributed systems, Kubernetes and observability experience.
  • Ability to design database architecture, runbooks, and incident procedures.

Responsibilities

  • Design, deploy, and own highly available Vitess and Cassandra DB architectures.
  • Document runbooks, escalation guides, and SOPs for SRE DB Engineers.
  • Optimize DB performance via tuning, indexing, and capacity planning.
  • Own backup, recovery, replication, and disaster recovery strategies.
  • Lead DR testing and incident response for database systems.
  • Drive security hardening, access control, patching, and compliance.
  • Collaborate on resharding, replication, and architecture decisions.
  • Support on-call rotations and post-incident reviews.
  • Develop automation for routine DB tasks and tooling pipelines.
  • Integrate runbooks with monitoring, CI/CD, and IaC frameworks.
  • Mentor Level 1/2 DB Engineers and create training materials.
  • Identify automation opportunities and operational efficiencies.

Skills

Linux
Automation
Distributed systems
Kubernetes
Observability
Incident response
Database design

Tools

Prometheus
Grafana
ELK
FireHydrant
Terraform
Ansible
Jenkins
kubectl
Vitess
mysqlsh
Docker

Job description

Remote Full-time Not specified Product

Job Description

About Backblaze

Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands.

Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $100m in revenue and is the leading specialized storage cloud – managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals.

But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking aSite Reliability Engineer III (DBA)to join our team!

About the Role:

Individuals fulfilling this role will be responsible for ensuring the stability, scalability, and reliability of our production database systems, primarily Vitess (distributed MySQL) and Cassandra, alongside the rest of our production services and infrastructure. This role carries the same on-call, incident response, and service ownership expectations as other SRE IIIs, with database systems serving as the area of deepest technical ownership.

Because our SRE Database Engineering function is new, this role will also help establish its operational foundation by designing database architecture and developing the runbooks, escalation guidance, procedures, and training materials that our Level 1 and Level 2 SRE Database Engineers will use as they onboard. The ideal candidate will have strong experience with production database systems, Linux, automation, distributed systems, Kubernetes, observability, and incident response, with a proactive approach to reliability and operational excellence.

What You’ll Do:

  • Design, deploy, and own highly available database architecture for Vitess (distributed MySQL) and Cassandra
  • Establish and document operational procedures, runbooks, and escalation guidance for Level 1 and Level 2 SRE Database Engineers
  • Optimize database performance through query tuning, indexing strategies, schema design, and capacity planning
  • Own database backup, recovery, replication, and disaster recovery strategies
  • Perform and validate disaster recovery testing and database recovery procedures
  • Drive database security, access control, patching, hardening, and compliance practices
  • Partner with the DBA and Data Infrastructure teams on resharding, capacity planning, replication, and architecture decisions for sharded MySQL environments
  • Service Reliability & Operations:
    • Support the availability and durability of critical services across production environments
    • Monitor service health using SLIs, SLOs, error budgets, monitoring, logging, and alerting platforms
    • Partner with service owners to define and improve SLIs, SLOs, error budget policies, and alerting
    • Participate in on-call rotations, incident response, root cause analysis, and post-incident reviews
    • Serve as an escalation point for complex database production incidents
    • Follow established ITIL/OSS processes including incident, change, problem, and capacity management
    • Take ownership of operational issues and drive projects from problem discovery through resolution
  • Automation & Tooling:
    • Develop automation for common operational and database administration tasks to reduce manual intervention and operational toil
    • Contribute to monitoring, logging, and alerting frameworks including Prometheus, Grafana, Catchpoint, and ELK
    • Help integrate operational runbooks and incident response workflows with FireHydrant
    • Work with CI/CD pipelines, configuration management, and infrastructure as code tools including Terraform, Ansible, and Jenkins
    • Develop scripts using Bash, Python, Go, or similar technologies to improve reliability and operational efficiency
    • Operate and troubleshoot containerized production environments using Kubernetes and Docker
    • Work within Kubernetes and Vitess environments using technologies such as kubectl, mysqlsh, and Vitess keyspaces
  • Project Management:
    • Lead Production Readiness Reviews (PRRs) for functionality being handed off from engineering partner teams
    • Support the operational readiness of new database-backed services before they enter production
    • Build training plans, onboarding materials, and technical documentation for new Level 1 and Level 2 SRE Database Engineers
    • Partner with Engineering, Product, Operations, and DBA/Data Infrastructure teams on reliability initiatives
    • Assist with capacity planning, disaster recovery exercises, database migrations, and infrastructure projects
    • Work with vendors and service providers to troubleshoot service issues and track SLA performance
    • Identify opportunities for automation and process efficiency
  • Respond to and resolve production database, infrastructure, and service incidents
  • Troubleshoot and elevate database, Linux, networking, application, and infrastructure issues as needed
  • Participate in the on-call rotation and serve as an escalation point for database-related incidents
  • Lead or contribute to root cause analysis and post-incident reviews
  • Identify recurring issues and develop long-term corrective actions to improve reliability
  • What we value:
    • A proactive mindset with a can-do attitude
    • Someone who can work independently, take ownership, and drive complex technical problems through resolution
    • Someone who steps up, supports teammates, mentors others, and shares knowledge freely
    • Strong problem-solving skills and a willingness to learn new technologies
    • Curiosity, reliability, and a desire to improve the reliability and scalability of production systems
  • ESPP program

To provide greater transparency to candidates, we share base pay ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar-stage growth companies. Final offer amounts are determined by multiple factors, including candidate location, skills, depth of work experience, and relevant licenses/credentials, and may vary from the amounts listed below.

The expected salary range for this role is – $125,000 – $150,000.

At Backblaze, we value being fair and good to our customers, partners, and employees. That’s why diversity, equity, and inclusion are at the core of our values. We are committed to fostering a workforce where all employees feel a sense of belonging regardless of race, ethnicity, nationality, gender, sexual orientation, age, religion, socio-economic status, ability, veteran status, and education. We believe that our dedication to cultivating a diverse workspace not only allows us to better serve our customers in over 175 countries but further reinforces our commitment to doing the right thing.We are proud to be an Equal Opportunity Employer.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer III (DBA)
Site Reliability Engineer III (DBA)

Backblaze External Website • United States

Remote
USD 125,000 - 150,000
MacBook Pro for work
Stipend for workstation setup
Flexible vacation policy
+6
Site Reliability Engineer ll (DBA)
Site Reliability Engineer ll (DBA)

Backblaze • United States

Remote
ARS 3,000,000 - 5,000,000
Site Reliability Engineer II
Site Reliability Engineer II

Webhosting • Northern (KY)

On-site
USD 110,000 - 160,000
Data Center Technician I
Data Center Technician I

Webhosting • Reston (VA), Northern (KY)

On-site
USD 62,000 - 72,000
ESPP program
Data Center Technician I
Data Center Technician I

Backblaze • Chandler (AZ)

On-site
USD 62,000 - 72,000
Healthcare for family
RSU grants
ESPP program
+6
Data Center Technician I (On-site in Chandler, AZ)
Data Center Technician I (On-site in Chandler, AZ)

Backblaze External Website • Chandler (AZ)

On-site
USD 62,000 - 72,000
Healthcare for family
Dental & Vision
401K
+7
Data Center Technician I (On-site in Chandler, AZ)
Data Center Technician I (On-site in Chandler, AZ)

Webhosting • Chandler (AZ), Northern (KY)

On-site
USD 62,000 - 72,000
ESPP program
Senior Software Engineer
Senior Software Engineer

Webhosting • San Mateo (CA)

On-site
USD 180,000 - 240,000
ESPP program
Data Center Operations Manager
Data Center Operations Manager

Webhosting • Chandler (AZ)

On-site
USD 110,000 - 140,000
Director, Enterprise Systems & Architecture New Remote - US
Director, Enterprise Systems & Architecture New Remote - US

Backblaze • Northern (KY)

On-site
USD 190,000 - 220,000
Competitive compensation
Equity package
401(k) with matching
+5