Principal Platform Engineer - London

heidihealth.com.au

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health and dental cover
£700 annual L&D budget
£100/month health/wellness allowance
£500 home office budget
Parental leave: 26 weeks (primary) & 4
Equity

Job summary

Heidi is seeking an experienced Site Reliability Engineer for the core Platform/SRE team. You’ll own production reliability, participate in on-call incidents, and maintain daily operations for Heidi’s platform.

We welcome mid-level to senior SREs who enjoy hands-on ownership. You’ll work with Kubernetes, AWS, Terraform, and observability tools to reduce toil and keep systems healthy across our growing healthcare platform.

Qualifications

  • 3–6+ years in SRE/DevOps/Platform roles.
  • Experience supporting production systems and on-call rotations.
  • Comfortable debugging live systems under pressure.
  • Experience operating cloud infrastructure (AWS preferred).
  • Working knowledge of Kubernetes and containerised workloads.
  • Infrastructure as Code experience (Terraform).
  • Familiarity with monitoring and alerting tools (Datadog, Prometheus).
  • Scripting or automation experience (Python, Bash).
  • Passion for AI, hands-on building or side projects.
  • Prefers building over requesting.
  • Confident in thinking; open to being wrong.

Responsibilities

  • Participate in on-call and incident response, lead incidents end-to-end over time.
  • Improve operational reliability through alerts, automation and system changes.
  • Own parts of the production environment: Kubernetes clusters and core services.
  • Strengthen observability with dashboards, logs and traces.
  • Reduce operational toil via automation and better tooling.
  • Support safe change with improved deployments and runbooks.
  • Contribute to incident response processes with blameless post-mortems.
  • Collaborate with engineers to improve production readiness and reliability.

Skills

SRE/DevOps
AWS
Kubernetes
Terraform
Monitoring/Alerting
Scripting (Python/Bash)
On-call
Incident response
Automation mindset
AI interest

Tools

Datadog
Prometheus

Job description

We’re Heidi.

We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. Our platform has absorbed the administrative chaos of 175 million patient visits and supported 67 million clinical hours. Today, we support 2.8 million patient sessions a week across 190+ countries, in 110 languages and over 200 specialties.

Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s capacity for care and strengthen the human connection at its heart.

We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again.

We’ve grown annual recurring revenue from $1 million to $50 million in two years.

To go further, we’ve secured US$340 million: a $100 million Series C led by Blackbird, with Phoenix Court, Point72 Private Investments and Headline, alongside a $240 million growth investment led by General Catalyst’s Customer Value Fund.

If you want to join us in doing work worth shipping, jump in.

The role

This role sits in the core Platform/SRE team that owns production. You’ll work directly on incident response, on-call, system reliability, and day-to-day operations for Heidi’s platform.

We’re open to candidates who are strong mid-level SREs ready to take on more ownership, as well as senior SREs who enjoy being hands-on in operations. The role is intentionally ops-heavy and focused on keeping real systems healthy in production.

What you’ll do
  • Participate in on-call and incident response: Respond to production incidents, contribute to service restoration, and support clear communication during incidents. Over time, take increasing responsibility for leading incidents end-to-end.

  • Improve operational reliability: Identify recurring issues and reliability risks, and drive fixes through better alerting, automation, system changes, or process improvements.

  • Own parts of the production environment: Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services, with growing ownership as familiarity increases.

  • Strengthen observability: Improve dashboards, alerts, logs, and traces so issues are detected earlier and diagnosed faster, with a strong focus on actionable signals.

  • Reduce operational toil: Automate repetitive tasks, simplify runbooks, and improve tooling to make on-call and day-to-day operations easier and safer.

  • Support safe change: Improve deployments, rollback mechanisms, and operational readiness to reduce the risk of incidents caused by change.

  • Contribute to operational practices: Write and maintain runbooks, participate in blameless post-mortems, and help improve incident response processes over time.

  • Collaborate closely with engineers: Work with product and feature teams to improve production readiness, service ownership, and reliability expectations.

What you'll need
  • 3–6+ years in SRE, DevOps, Platform, or operations-heavy engineering roles.

  • Experience supporting production systems and participating in on-call rotations.

  • Comfortable debugging live systems under pressure.

  • Experience operating cloud infrastructure (AWS preferred).

  • Working knowledge of Kubernetes and containerised workloads.

  • Infrastructure as Code experience (Terraform or similar).

  • Familiarity with monitoring and alerting tools (Datadog, Prometheus, etc).

  • Scripting or automation experience (Python, Bash, or similar).

  • Passion for AI, shown through hands-on building, prototyping, or side projects.

  • You default to building over requesting.

  • You're confident in your thinking and open to being wrong. Great ideas win regardless of who surfaces them.

How we show up
  • Build for the next decade, not next quarter. Our targets are outrageous on purpose. The world's health doesn't have the luxury of incrementalism.

  • Lead, don't wait. We treat tomorrow's problems today. Sometimes we build what's needed before it's wanted, and we're fine with that.

  • Follow the evidence. Trust the patient. We pursue truth relentlessly. But when the subjective and objective disagree, we treat the patient, not the numbers. Ego is a comorbidity we can't afford.

  • Own the outcome. Everyone here carries the company. Raise problems with solutions, solve them end-to-end, and never be a bystander.

  • Ship, measure, go again. A button today, a workflow tomorrow. More iterations beat better planning. We're precise at pace, not reckless.

  • Live in clinicians' reality. Not the ideal workflow, the twenty-patients-before-lunch actual one. We build for exhausted humans, and we'd better be decent ones while we do it.

Why Heidi?

You’ll join a team measured on real-world impact, not press cycles. We live and breathe the challenges of modern health systems, and are laser-focused on exacting the change we’d like to see. We’re medicos, engineers, builders and designers who’ve felt (on every side of the equation) what non-care feels like, the moral and practical toll as a provider or receiver.

Building what we’re building isn’t always easy. Standard tech playbooks routinely fail in healthcare, and the friction of the medical system will test your perseverance. But we didn’t choose easy, we chose a purpose-driven mission that actually matters. We hold ourselves to a higher standard because healthcare demands it. If you join Heidi, you recognize that the deeper question isn’t whether AI can solve the global healthcare crisis, but whose hands will shape it.

True A-players progress extremely fast here. The nature of the scale‑up game is demanding, but we value sustainable performance and mental health. You're trusted to perform, and you set your schedule. We operate on outcomes > inputs, not process theatre. We all take the bins out, metaphorically and literally.

We take care of you.

We offer health and dental cover, a £700 annual learning and development budget, a £100/month health and wellness allowance, a £500 home office budget, 26 weeks paid primary parental leave and 18 weeks paid secondary parental leave, fertility support up to £7,000, four weeks of work from anywhere per year, and serious equity.

We chose to open-source our benefits hub, if you care to take a peek.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Platform Engineer - London
Principal Platform Engineer - London

black.ai • Greater London

On-site
GBP 70,000 - 110,000
Equity from day one
Private medical and dental cover
Home office budget
Principal Platform Engineer - London
Principal Platform Engineer - London

Heidi Health Corp. • Greater London

Hybrid
GBP 90,000 - 130,000
Equity from day one
Private medical and dental cover (Bupa
Home office budget
+2
Customer Support Engineer - London
Customer Support Engineer - London

Blackbird Ventures Pty • Greater London

On-site
GBP 42,000 - 65,000
Health & dental cover
£700 learning & development budget
£100/month health & wellness allowance
+4
Customer Support Engineer - London
Customer Support Engineer - London

Heidi • Greater London

Hybrid
GBP 40,000 - 70,000
Health and dental cover
£700 annual learning & development
£100/month health & wellness allowance
+6
Customer Support Engineer - London
Customer Support Engineer - London

heidihealth.com.au • Greater London

On-site
GBP 60,000 - 75,000
Health and dental cover
Learning and development budget
Wellness allowance
+2
Customer Support Engineering Lead - London
Customer Support Engineering Lead - London

Heidi • Greater London

On-site
GBP 90,000 - 120,000
Health & dental
Learning budget £700
Wellness allowance £100/m
+5
Customer Support Engineering Lead - London
Customer Support Engineering Lead - London

Blackbird Ventures Pty • Greater London

On-site
GBP 90,000 - 120,000
Health and dental cover
Learning budget (£700)
Wellness allowance (£100/month)
+4
Customer Support Engineering Lead - London
Customer Support Engineering Lead - London

Heidi Health Corp. • Greater London

Hybrid
GBP 90,000 - 130,000
Health and dental
Learning budget
Wellness allowance
+5
Customer Support Engineering Lead - London
Customer Support Engineering Lead - London

heidihealth.com.au • Greater London

Hybrid
GBP 90,000 - 130,000
Health and dental cover
£700 learning budget
£100/month wellness allowance
+6
Customer Support Engineer - London
Customer Support Engineer - London

Heidi Health Corp. • Greater London

Hybrid
GBP 45,000 - 70,000
Health and dental cover
£700 learning budget
£100 monthly health/wellness allowance
+6