Lead SRE & Platform Reliability

Heidi Health Corp.

City of Melbourne

On-site

AUD 140,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Learning budget ($1k)
Health/wellness allowance
Home office budget
Parental leave 26 weeks
Parental leave 2nd 18 weeks
Fertility support
Work from anywhere 4 weeks
Equity

Job summary

Heidi Health Corp. is expanding its core Platform/SRE team. You will lead a small SRE group, stay hands-on with incidents, and scale the team as Heidi grows. You’ll shape reliability standards, hire, and structure the team while remaining deeply involved in day‑to‑day operations.

You’ll own production, improve dashboards and alerts, and collaborate with engineers to ensure production readiness. This is a hands‑on leadership role requiring strong incident response and cloud/Kubernetes expertise.

Qualifications

  • 7+ years in SRE, DevOps, platform, or operations-heavy roles, with team leadership experience.
  • Proven record of hiring, coaching, and growing engineers.
  • Experience supporting production systems and on‑call rotations.
  • Ability to debug live systems under pressure.
  • Strong cloud infra experience at scale (AWS preferred).
  • Hands-on with Kubernetes and containerized workloads.
  • Experience with infrastructure as code (Terraform or similar).
  • Proficient with monitoring/alerting tools (Datadog/Prometheus) and alerting strategies.
  • Scripting/automation experience (Python, Bash).
  • Experience defining owning SLOs, error budgets, and capacity planning.

Responsibilities

  • Participate in on‑call and incident response; lead incidents end‑to‑end.
  • Improve operational reliability via alerts, automation, and process improvements.
  • Own the production environment across Kubernetes and cloud infra.
  • Strengthen observability with dashboards, logs, and traces.
  • Reduce operational toil through tooling and runbooks.
  • Support safe change with deployments, rollbacks, and readiness checks.
  • Write and maintain runbooks; conduct blameless post‑mortems.
  • Collaborate with engineers on production readiness and service ownership.
  • Lead and grow the SRE team; manage headcount and on‑call structure.
  • Decide team focus and balance reliability work with capacity planning.

Skills

SRE leadership
AWS cloud
Kubernetes
IaC (Terraform)
Monitoring (Datadog/Prometheus)
Automation scripting (Python/Bash)
Incident response
SLOs & capacity planning

Tools

Terraform
AWS
Datadog
Prometheus
Kubernetes

Job description

Heidi Health Corp. is expanding its core Platform/SRE team. You will lead a small SRE group, stay hands-on with incidents, and scale the team as Heidi grows. You’ll shape reliability standards, hire, and structure the team while remaining deeply involved in day‑to‑day operations.

You’ll own production, improve dashboards and alerts, and collaborate with engineers to ensure production readiness. This is a hands‑on leadership role requiring strong incident response and cloud/Kubernetes expertise.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Lead, Developer Platform & Reliability
Engineering Lead, Developer Platform & Reliability

heidihealth.com.au • City of Melbourne

On-site
AUD 180,000 - 250,000
L&D budget
Health allowance
Home office budget
+4
Engineering Lead - Global Cloud Reliability
Engineering Lead - Global Cloud Reliability

Heidi • City of Melbourne

On-site
AUD 180,000 - 240,000
Learning and development budget
Health & wellness allowance
Home office budget
+4
Cloud & Reliability Engineering Lead
Cloud & Reliability Engineering Lead

Heidi • Sydney

On-site
AUD 140,000 - 200,000
L&D budget
Health & wellness allowance
Home office budget
+4
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Heidi Health Corp. • City of Melbourne

On-site
AUD 140,000 - 210,000
Learning budget ($1k)
Health/wellness allowance
Home office budget
+5
Platform & Reliability Engineering Lead
Platform & Reliability Engineering Lead

black.ai • City of Melbourne

On-site
AUD 180,000 - 250,000
Learning & development budget
Health and wellness allowance
Home office budget
+3
Senior SRE Lead: Multi-Cloud Platform Reliability
Senior SRE Lead: Multi-Cloud Platform Reliability

Leidos Australia • City of Knox

On-site
AUD 130,000 - 160,000
Family Friendly workplace
Diversity and inclusion initiatives
Engineering Lead - Reliability & Cloud
Engineering Lead - Reliability & Cloud

Heidi • City of Melbourne

On-site
AUD 180,000 - 240,000
Learning and development budget
Health & wellness allowance
Home office budget
+4
Senior SRE & Platform Lead – Multi-Cloud & Secure Ops
Senior SRE & Platform Lead – Multi-Cloud & Secure Ops

Leidos • City of Melbourne

On-site
AUD 180,000 - 240,000
Remote SRE - Core Cloud Reliability & Incident Response
Remote SRE - Core Cloud Reliability & Incident Response

Opentalent • City of Brisbane

Hybrid
AUD 140,000 - 180,000
Senior SRE — Cloud, OpenShift & Platform Reliability
Senior SRE — Cloud, OpenShift & Platform Reliability

Egis Group • Victoria

On-site
AUD 120,000 - 180,000