Staff Database Reliability Engineer

HRB

Canada

Presencial

CAD 120.000 - 190.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Convierte este puesto en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

RRSP matching

Descripción de la vacante

HRB is seeking an accomplished Database Reliability Engineer to design and evolve globally distributed database systems, balancing performance with reliability across bare-metal and cloud environments. You will lead complex initiatives, reduce operational complexity, and drive automation for provisioning, deployment, and change management.

You will collaborate with platform engineering, infrastructure, software, operations, and data teams to improve reliability, latency, and scalability while

Formación

  • 7+ years of experience spanning database reliability engineering, database administration, site reliability engineering focused on data systems, platform engineering, or a comparable infrastructure role.
  • Experience operating large-scale production database systems where reliability, latency, availability, and performance matter.
  • Deep expertise with relational and distributed data systems and production experience with technologies such as MySQL, MariaDB, Galera, PostgreSQL, Kafka, Aerospike, Redis, StarRocks, or Vertica.
  • Experience deploying and operating stateful systems on Kubernetes and bare-metal infrastructure.
  • Strong automation instincts and experience with Terraform, Ansible, Helm, GitLab CI/CD, ArgoCD, or similar tools.
  • Ability to automate infrastructure and operational workflows using languages such as Python, Go, Bash/Shell, Java, or Kotlin.
  • Experience building or operating observability systems (Prometheus, Grafana, Loki, etc.).
  • Understanding of database change management with tools like Liquibase, Flyway, Alembic, or similar.
  • Ability to troubleshoot complex failures across distributed infrastructure and derive root causes.
  • Comfort operating at staff level, setting direction and mentoring other engineers.
  • Excellent communication of complex decisions to technical and non-technical stakeholders.

Responsabilidades

  • Design and evolve globally distributed database systems for performance, availability, scalability, and continuity in both bare-metal and cloud environments.
  • Lead complex initiatives, establish technical direction, and influence engineering decisions across teams without direct reporting authority.
  • Identify opportunities to reduce operational complexity and consolidate database and infrastructure footprints.
  • Design and implement automation for provisioning, configuration, deployment, upgrades, testing, maintenance, and change management use cases.
  • Establish durable practices for operating critical data infrastructure, including monitoring, incident response, capacity planning, security, IAM, auditing, and traceability.
  • Diagnose performance and reliability problems across latency-sensitive infrastructure and drive root-cause resolution.
  • Lead retrospectives and root-cause analyses for significant data-system incidents and translate findings into tooling and architectural improvements.
  • Evaluate emerging technologies and develop roadmaps for the data platform; communicate priorities to engineering and business stakeholders.
  • Mentor engineers via architecture reviews, guidance, documentation, and hands-on collaboration.

Conocimientos

Database reliability engineering
Technical leadership
Staff-level influence
Distributed data systems
Root cause analysis
Automation
Observability
Communication to stakeholders
Mentorship

Herramientas

Terraform
Ansible
Helm
GitLab CI/CD
ArgoCD
Python
Go
Bash/Shell
Java
Kotlin

Descripción del empleo

Our client builds a high throughput system powering hundreds of billions of daily transactions, each completed within milliseconds, across globally distributed infrastructure designed for reliability and efficiency. They have been quietly bootstrapping and growing in line with revenue for over 2 decades. They maintain independence from external stakeholders, enabling them to chart their course, maintain a long‑term perspective and build an enduring, sustainable business that currently employs ~600 team members globally.


Data is central to nearly every part of their platform. High‑volume event and transactional data powers customer reporting, billing, marketplace analytics, experimentation, machine‑learning workflows, and product experiences. The organization is evolving from centralized, batch‑oriented reporting toward a platform‑driven architecture combining batch, streaming, and asynchronous processing, with APIs becoming a primary interface to data.


Our client has been stubbornly racking and stacking infrastructure around the world for the duration of their existence, a habit that allows them to price themselves at the cost of electricity while their competitors are mired in rising cloud infrastructure costs. This long‑term, somewhat contrarian, thinking puts them in a position to offer stability and career longevity evidenced by robust benefits that include RRSP matching.


This is an IC role for someone who can move comfortably between deep technical problem solving and broader architectural leadership. You’ll help define how database systems are designed, operated, automated, observed, and evolved across the organization. You’ll work closely with platform engineering, infrastructure, software engineering, operations, and data teams to improve reliability and scalability while championing simplicity in a sophisticated technology landscape.


Key Responsibilities


  • Database architecture & reliability: Design and evolve globally distributed database systems for performance, availability, scalability, and continuity in both bare‑metal and cloud environments

  • Technical leadership: Lead complex initiatives, establish technical direction, and influence engineering decisions across teams without relying on direct reporting authority

  • Platform simplification: Identify opportunities to reduce operational complexity and consolidate the database and infrastructure technology footprint

  • Automation: Design and implement automation for provisioning, configuration, deployment, upgrades, testing, maintenance, and database change management use cases

  • Operational excellence: Establish durable practices for operating critical data infrastructure, including monitoring, incident response, capacity planning, maintenance, security, IAM, auditing, and traceability

  • Performance engineering: Diagnose difficult performance and reliability problems across globally distributed, latency‑sensitive infrastructure and drive issues through root cause to durable resolution

  • Incident learning: Lead retrospectives and root‑cause analysis for significant data‑system incidents and convert findings into improvements to tooling, architecture, and operational practices

  • Technical roadmap: Evaluate emerging technologies, develop roadmaps for the data platform, and communicate priorities and trade‑offs across engineering and business stakeholders

  • Mentorship: Raise the technical bar for engineers working with data infrastructure through architecture reviews, technical guidance, documentation, and hands‑on collaboration


Your Know-How


  • You have 7+ years of experience spanning database reliability engineering, database administration, site reliability engineering focused on data systems, platform engineering, or a comparable infrastructure role

  • You have operated large‑scale production database systems where reliability, latency, availability, and performance genuinely matter

  • You bring deep expertise with relational and distributed data systems and meaningful production experience with several technologies (such as MySQL, MariaDB, Galera, PostgreSQL, Kafka, Aerospike, Redis, StarRocks, or Vertica)

  • You have experience deploying and operating stateful systems across Kubernetes and bare‑metal infrastructure

  • You understand relational data architecture and can make informed decisions around schema design, replication, availability, performance, scale, and continuity

  • You have strong automation instincts and experience with infrastructure tooling such as Terraform, Ansible, Helm, GitLab CI/CD, ArgoCD, or comparable technologies

  • You can automate infrastructure and operational workflows using Python, Go, Bash/Shell, Java, Kotlin, or another appropriate programming language

  • You have experience building or operating effective observability systems (using technologies such as Prometheus, Grafana, Loki, or equivalent tooling)

  • You understand database change management and have worked with tooling such as Liquibase, Flyway, Alembic, or similar systems

  • You can troubleshoot complex failures across distributed infrastructure, reason from symptoms through multiple layers of a system, and identify root causes rather than treating symptoms

  • You’re comfortable operating at “staff level” (aka setting direction, navigating ambiguity, mentoring other engineers, balancing immediate operational needs against long‑term architecture, and influencing teams outside your immediate domain)

  • You can communicate complex technical decisions clearly to both deeply technical colleagues and stakeholders without the same infrastructure background


It’s a bonus if


  • You have experience operating databases across globally distributed, ultra‑low‑latency infrastructure

  • You have meaningful experience with streaming and large‑scale data technologies such as Kafka, Spark, Flink, Hadoop, Trino/Presto, or Iceberg

  • You’ve operated infrastructure spanning both public cloud and privately operated data centres

  • You have experience simplifying or consolidating a large database technology footprint without compromising reliability or developer productivity

  • You have experience helping establish security, IAM, auditing, and traceability practices for critical data infrastructure

  • You’ve worked in adtech, financial infrastructure, high‑frequency systems, large‑scale marketplaces, or another environment where extremely high transaction volume and low latency are fundamental engineering constraints

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Backend Software Engineer (Online Storage)
Senior Backend Software Engineer (Online Storage)

Affirm • Ottawa

A distancia
CAD 110.000 - 150.000
Remote-first compensation structure
Generous time off
Health benefits
+4
Database Reliability Engineer
Database Reliability Engineer

Humankind Global Recruitment • Halifax

Presencial
CAD 100.000 - 140.000
Competitive compensation package
Medical & health benefits
RRSP matching
Principal ML Engineer
Principal ML Engineer

HRB • Canadá

A distancia
CAD 140.000 - 210.000
RRSP matching
Principal SW Engineer
Principal SW Engineer

HRB • Toronto, Kitchener

Presencial
CAD 140.000 - 200.000
RRSP matching
Chief Analyste, Financial Performance (Hybrid)
Chief Analyste, Financial Performance (Hybrid)

National Bank • Montreal (administrative region)

Presencial
CAD 120.000 - 180.000
RRSP matching
Data Platform Developer
Data Platform Developer

Computationaldesign • Quebec

Presencial
CAD 110.000 - 170.000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Bedford

Presencial
CAD 140.000 - 200.000
Software Engineer II
Software Engineer II

FiveTran • Toronto

Presencial
CAD 90.000 - 150.000
Software Engineer: Backend Systems
Software Engineer: Backend Systems

Aloe • Vancouver

Presencial
CAD 110.000 - 170.000
Competitive salary
Benefits
Career growth opportunities
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Markham

Presencial
CAD 120.000 - 180.000