Head of Platform Reliability

Cielo Projects

New York (NY)

On-site

USD 165,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Quant is building its New York team to recruit a Head of Platform Reliability who will own the reliability of our critical platform infrastructure. You will define service levels, establish the error budget policy, and have the authority to stop a release when the budget is spent.

You will lead the engineering team responsible for running it, designing the rotation and ensuring robust, predictable production systems.

Qualifications

  • Led site reliability or production engineering somewhere failure had consequences.
  • Ran error budgets in practice — used one to stop a release, and defended the call
  • Honest experience of continuous coverage — what burns people out, rotation you can live with
  • Kubernetes, Terraform, and an observability stack you have actually debugged, rather than configured.

Responsibilities

  • Own reliability by defining service levels, setting the error budget policy, and having authority to stop a release when the budget is spent.
  • Build and lead the engineering team that runs it, on a rotation you design.
  • Tackle stateful infrastructure where restarts aren’t a strategy and protocol behavior can look like infrastructure.
  • Collaborate across teams to ensure platform reliability for regulated money in production.

Skills

Site reliability
Production engineering
Error budgets
Kubernetes
Terraform

Tools

Observability stack

Job description

  • Compensation: USD 165000 - USD 210000 - yearly
Company Description

We're partnering withQuant, a global leader in digital transformation and technology solutions, seeking aHead of Platform Reliabilityto join their team.

About Quant:

Almost all the money in the economy is commercial bank money. Deposits, sitting on bank balance sheets, moving through payment systems designed decades before anyone had a reason to make money programmable. Almost none of it moves on chain.

Quant builds the infrastructure that changes that. Our technology lets a bank issue, move and settle its own money on programmable rails while staying connected to the systems it already runs on. That constraint is the reason this has taken as long as it has.

Central banks and commercial banks have built on our platform, including work on the digital pound and the digital euro. We are now building our team in New York.

Job Description

Head of Platform Reliability

Chain infrastructure carrying regulated money, in production, for institutions that hold it to the highest possible standard.

You would own its reliability. That means defining the service levels, setting the error budget policy, and holding the authority to stop a release when the budget is spent — a call that stands regardless of seniority.

You would build and lead the engineering team that runs it, on a rotation you design. We would rather you designed it properly than inherited something and patched it.

Chain platforms fail differently from the systems most reliability engineers have run. State is expensive, restarts are not a strategy, and a surprising amount of what looks like infrastructure turns out to be protocol behaviour. If that sounds interesting rather than daunting, we should talk.

Qualifications

You will need

  • To have led site reliability or production engineering somewhere failure had consequences.
  • To have run error budgets in practice — used one to stop a release, and defended the call
  • Honest experience of continuous coverage — what burns people out, what does not, and the difference between a rotation on paper and one people can live with
  • Kubernetes, Terraform, and an observability stack you have actually debugged, rather than configured.

Useful, not essential

  • Blockchain nodes in production.
  • Banking, payments, or somewhere else an auditor asked you to prove something.
Additional Information

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Reliability Technical Lead
Platform Reliability Technical Lead

Cielo Projects • New York (NY)

On-site
USD 170,000 - 220,000
Platform Reliability Lead for Regulated Money Systems
Platform Reliability Lead for Regulated Money Systems

Cielo Projects • New York (NY)

On-site
USD 165,000 - 210,000
Platform Engineer - Reliability
Platform Engineer - Reliability

Squarepoint • Houston (TX)

On-site
USD 100,000 - 130,000
Service Management Lead
Service Management Lead

Cielo Projects • New York (NY)

On-site
USD 130,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE Clear Europe Limited • Jacksonville (FL)

On-site
USD 120,000 - 170,000
Head of SRE - Quant Trading
Head of SRE - Quant Trading

Acquire Me • Chicago (IL)

On-site
USD 150,000 - 200,000
Senior Software Engineer Cloud Platform (Infrastructure, Due Diligence & Reliability)
Senior Software Engineer Cloud Platform (Infrastructure, Due Diligence & Reliability)

Outsource Talent Pro • New York (NY)

On-site
USD 140,000 - 160,000
Cutting-edge technology
Collaborative environment with experts
Impactful projects in quantum computing
Platform Reliability Lead: Production Systems Expert
Platform Reliability Lead: Production Systems Expert

Cielo Projects • New York (NY)

On-site
USD 170,000 - 220,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000