Lead Site Reliability Engineer

Heidi Health Corp.

City of Melbourne

On-site

AUD 140,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Learning budget ($1k)
Health/wellness allowance
Home office budget
Parental leave 26 weeks
Parental leave 2nd 18 weeks
Fertility support
Work from anywhere 4 weeks
Equity

Job summary

Heidi Health Corp. is expanding its core Platform/SRE team. You will lead a small SRE group, stay hands-on with incidents, and scale the team as Heidi grows. You’ll shape reliability standards, hire, and structure the team while remaining deeply involved in day‑to‑day operations.

You’ll own production, improve dashboards and alerts, and collaborate with engineers to ensure production readiness. This is a hands‑on leadership role requiring strong incident response and cloud/Kubernetes expertise.

Qualifications

  • 7+ years in SRE, DevOps, platform, or operations-heavy roles, with team leadership experience.
  • Proven record of hiring, coaching, and growing engineers.
  • Experience supporting production systems and on‑call rotations.
  • Ability to debug live systems under pressure.
  • Strong cloud infra experience at scale (AWS preferred).
  • Hands-on with Kubernetes and containerized workloads.
  • Experience with infrastructure as code (Terraform or similar).
  • Proficient with monitoring/alerting tools (Datadog/Prometheus) and alerting strategies.
  • Scripting/automation experience (Python, Bash).
  • Experience defining owning SLOs, error budgets, and capacity planning.

Responsibilities

  • Participate in on‑call and incident response; lead incidents end‑to‑end.
  • Improve operational reliability via alerts, automation, and process improvements.
  • Own the production environment across Kubernetes and cloud infra.
  • Strengthen observability with dashboards, logs, and traces.
  • Reduce operational toil through tooling and runbooks.
  • Support safe change with deployments, rollbacks, and readiness checks.
  • Write and maintain runbooks; conduct blameless post‑mortems.
  • Collaborate with engineers on production readiness and service ownership.
  • Lead and grow the SRE team; manage headcount and on‑call structure.
  • Decide team focus and balance reliability work with capacity planning.

Skills

SRE leadership
AWS cloud
Kubernetes
IaC (Terraform)
Monitoring (Datadog/Prometheus)
Automation scripting (Python/Bash)
Incident response
SLOs & capacity planning

Tools

Terraform
AWS
Datadog
Prometheus
Kubernetes

Job description

We’re Heidi.

We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. In just 18 months, our clinical AI products have absorbed the administrative chaos of 73 million patient visits. Today, we support over 2.5 million patient sessions a week across 190+ countries.

Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s healthcare capacity and strengthen the human connection at its heart.

We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again.

If you don’t choose easy and you want to build something way bigger than yourself then, choose the challenge, choose Heidi.

The role

This role sits in the core Platform/SRE team that owns production. You'll lead a small SRE team today, growing it as Heidi scales, while staying hands-on in incident response, on-call, system reliability, and day-to-day operations. We're looking for someone who has already built or scaled a reliability team, not someone stepping into management for the first time. You'll set the standard for how the team operates, hire and structure it as headcount grows, and represent SRE in conversations with engineering leadership. The role stays ops-heavy: you're expected to be in the systems yourself, not just running a roadmap from a distance.

What you’ll do
  • Participate in on-call and incident response. Respond to production incidents, contribute to service restoration, and keep communication clear while things are on fire. Lead incidents end-to-end, including the ones outside your immediate team.
  • Improve operational reliability. Spot recurring issues and reliability risks, then drive fixes through better alerting, automation, system changes, or process improvements.
  • Own the production environment. Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services across the team's remit.
  • Strengthen observability. Build dashboards, alerts, logs, and traces that surface issues earlier and cut diagnosis time, with a bias toward signals people can actually act on.
  • Reduce operational toil. Automate the repetitive stuff, simplify runbooks, and improve tooling so on-call and daily operations get easier and safer over time.
  • Support safe change. Improve deployments, rollback mechanisms, and operational readiness so shipping changes doesn't mean rolling the dice on an incident.
  • Contribute to operational practices. Write and maintain runbooks, run blameless post-mortems, and raise the bar on incident response as the team learns.
  • Collaborate closely with engineers. Partner with product and feature teams on production readiness, service ownership, and what "reliable" should mean for their systems.
  • Lead and grow the SRE team. Manage the team you inherit day to day, then hire and onboard as Heidi scales. Own on-call structure, career development, and team norms as headcount grows.
  • Shape the team's direction. Decide where the team invests next, balance reliability work against team capacity, and represent SRE in planning with engineering leadership.
What you'll need
  • 7+ years in SRE, DevOps, platform, or operations-heavy engineering roles, including experience formally leading or managing a team.
  • A track record of hiring, coaching, and growing engineers, not just leading incidents solo.
  • Deep experience supporting production systems, including on-call rotations, and the credibility to still jump into an incident yourself.
  • You're comfortable debugging live systems under pressure.
  • Strong experience operating cloud infrastructure at scale (AWS preferred).
  • Solid hands-on experience with Kubernetes and containerized workloads in production.
  • Infrastructure as code experience (Terraform or similar).
  • Hands-on experience with monitoring and alerting tools such as Datadog or Prometheus, including designing alerting strategy, not just consuming dashboards.
  • Scripting or automation experience in Python, Bash, or similar.
  • Hands-on experience defining and owning SLOs, error budgets, and capacity planning, not just familiarity with the concepts.
Nice to have
  • Experience scaling a team through fast headcount growth, not just steady-state management.
  • Experience in regulated or security-sensitive environments.
  • Familiarity with databases, queues, and caches in production.
How we show up
  • Build for the next decade, not next quarter. Our targets are outrageous on purpose. The world's health doesn't have the luxury of incrementalism.
  • Lead, don't wait. We treat tomorrow's problems today. Sometimes we build what's needed before it's wanted, and we're fine with that.
  • Follow the evidence. Trust the patient. We pursue truth relentlessly. But when the subjective and objective disagree, we treat the patient, not the numbers. Ego is a comorbidity we can't afford.
  • Own the outcome. Everyone here carries the company. Raise problems with solutions, solve them end-to-end, and never be a bystander.
  • Ship, measure, go again. A button today, a workflow tomorrow. More iterations beat better planning. We're precise at pace, not reckless.
  • Live in clinicians' reality. Not the ideal workflow, the twenty-patients-before-lunch actual one. We build for exhausted humans, and we'd better be decent ones while we do it.
Why Heidi?

You’ll join a team focused on real-world impact over imaginary valuations and glossy PR. We live and breathe the challenges of modern health systems, and are laser-focused on exacting the change we’d like to see. We’re medicos, engineers, builders, and designers who’ve felt the moral and practical toll of what non-care feels like. True A-players progress extremely fast here. The nature of the scale-up game is demanding, but we value sustainable performance and mental health. You're trusted to perform, and you set your schedule. We operate on outcomes > inputs, not process theatre. We all take the bins out, metaphorically and literally.

Building what we’re building isn’t always easy. But we didn’t choose easy, we chose to build something that actually matters. We hold ourselves to a higher standard because healthcare demands it. If you join Heidi, you recognise that the deeper question isn’t whether AI can solve the global healthcare crisis, but whose hands will shape it. The work is hard, but you will trust and admire the people you work beside, and rest easy knowing you’re doing the defining work of your career.

We take care of you.
  • a $1,000 annual learning and development budget
  • a $150/month health and wellness allowance
  • a $500 home office budget
  • 26 weeks paid primary parental leave
  • 18 weeks paid secondary parental leave
  • fertility support up to $10,000
  • four weeks of work from anywhere per year
  • serious equity
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Lead - Reliability & Cloud
Engineering Lead - Reliability & Cloud

Heidi • City of Melbourne

On-site
AUD 180,000 - 240,000
Learning and development budget
Health & wellness allowance
Home office budget
+4
Engineering Lead, Cloud & Reliability (SRE)
Engineering Lead, Cloud & Reliability (SRE)

Heidi • Sydney

On-site
AUD 180,000 - 260,000
Learning & development budget
Health & wellness allowance
Home office budget
+3
Engineering Lead - Reliability & Cloud
Engineering Lead - Reliability & Cloud

Heidi • Sydney

On-site
AUD 180,000 - 260,000
Learning & development budget
Health & wellness allowance
Home office budget
+3
Engineering Lead - Reliability & Cloud
Engineering Lead - Reliability & Cloud

Heidi • Sydney

On-site
AUD 140,000 - 200,000
L&D budget
Health & wellness allowance
Home office budget
+4
Engineering Lead - Developer Platform
Engineering Lead - Developer Platform

heidihealth.com.au • City of Melbourne

On-site
AUD 180,000 - 250,000
L&D budget
Health allowance
Home office budget
+4
Engineering Lead - Developer Platform
Engineering Lead - Developer Platform

black.ai • City of Melbourne

On-site
AUD 180,000 - 250,000
Learning & development budget
Health and wellness allowance
Home office budget
+3
Head Of Implementation
Head Of Implementation

black.ai • City of Melbourne

On-site
AUD 180,000 - 260,000
Annual learning & development budget
Health and wellness allowance
Home office budget
+4
Support Engineer
Support Engineer

Heidi • City of Melbourne, Sydney

On-site
AUD 90,000 - 130,000
Learning budget
Health allowance
Home office budget
+5
Support Engineer
Support Engineer

black.ai • City of Melbourne

On-site
AUD 90,000 - 120,000
Learning budget
Health allowance
Home office budget
+4
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

heidihealth.com.au • City of Melbourne

On-site
AUD 180,000 - 230,000
Learning budget
Health/wellness allowance
Home office budget
+4