Senior Production Engineer

clearstreet

Greater London

On-site

GBP 110,000 - 140,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Clear Street in London is seeking a Production Engineer who owns the health and resilience of our production systems. You'll partner with engineering, operations, and business teams to translate daily pain points into automated solutions.

Half your time is spent supporting production and incidents; the other half building automation, tooling, and observability platforms to make on-call shifts easier and outages rarer.

Qualifications

  • Strong hands-on Python for automation and tooling.
  • Experience in SRE/Production or Platform Engineering with production ownership.
  • Proven automation and diagnostic tooling to reduce MTTR.

Responsibilities

  • Design and build monitoring and observability platforms to surface signals and reduce alert fatigue.
  • Develop automation and self-healing capabilities to accelerate recovery.
  • Analyze incidents to identify systemic trends and prevent recurring failures.
  • Create reusable runbooks and tooling to scale platform operations.
  • Collaborate with Platform Engineering to influence CI/CD safety and resilience.
  • Champion IaC and GitOps practices across the team.

Skills

Python
SRE/Production Eng
Automation tooling
Cloud-native tech
Kubernetes
Observability (Datadog)
Terraform
GitOps (ArgoCD, GA)
Distributed systems
Postgres/DBs
Communication
Problem solving

Tools

Datadog
Terraform
ArgoCD
GitHub Actions
Kafka
Postgres

Job description

About Clear Street:

Clear Street's mission is to give every sophisticated investor access to every asset, in every market, through a unified platform built for speed, transparency and scale.

We give our clients the technology, tools, and service once reserved for the largest institutions, rebuilt with modern infrastructure. Our single, cloud-native, end-to-end capital markets platform powers investor growth today and is transforming how they can interact with markets tomorrow.

For more information, visit https://clearstreet.io .

The Role

As a Production Engineer, you sit at the intersection of software reliability and operational excellence. You own the health, resilience, and recovery of our production systems-while spending equal energy innovating solutions that eliminate human toil, reduce incident blast radius, and raise the reliability bar across the entire platform. You will partner closely with engineering, operations, and business teams to understand daily pain points and translate them into lasting automated solutions. Half your time is spent in the trenches-supporting production, responding to incidents, and deeply understanding how our systems behave under real conditions. The other half is yours to build: automation, tooling, and observability platforms that make tomorrow's on-call shift meaningfully easier than today's.

You will work on challenges like:
  • Design and build comprehensive monitoring and observability platforms that surface the right signal at the right time-eliminating alert fatigue and accelerating root-cause analysis.
  • Develop intelligent automation and self-healing capabilities that diagnose issues, trigger recovery workflows, and reduce mean time to recovery (MTTR) without manual intervention.
  • Analyze incidents, identify systemic trends, and engineer solutions that prevent entire classes of failures from recurring.
  • Build reusable runbooks, diagnostic tooling, and recovery playbooks that turn tribal knowledge into scalable platform capabilities.
  • Create golden-path operational workflows-making the safest, most reliable path also the easiest one for engineering teams to follow.
  • Partner with Platform Engineering to influence CI/CD pipelines, deployment safety, and infrastructure resilience from a production reliability perspective.
  • Champion Infrastructure as Code, GitOps, and SRE best practices while helping teams adopt modern engineering workflows.
  • Continuously measure production health through SLIs/SLOs/SLAs, and drive engineering priorities based on reliability data.
  • Explore emerging technologies-including AI-assisted diagnostics and developer tooling-that transform how we operate production systems.
The Team

We believe resilient systems are built by engineers who understand them end to end. Our Production Engineering team is the first and last line of defense for our production platform. We treat reliability as a product, with uptime and engineer experience as our north stars. We combine the discipline of SRE with a builder's mindset: when we see a recurring problem, we build a solution-not a workaround.

If you're passionate about making production systems invisible to end users-and you get energy from both firefighting and building the systems that make fires less likely-you'll thrive here.

What We're Looking For

We're looking for engineers who combine operational instinct with a builder's discipline.

You should have:

  • Strong hands-on Python skills-this is your primary language for automation and tooling.
  • Experience in SRE, Production Engineering, Platform Engineering, or a related discipline with direct production ownership.
  • Proven track record of building automation and diagnostic tooling that improved recovery times or reduced operational toil.
  • Deep familiarity with cloud-native technologies-Kubernetes, containers, distributed systems-and how they fail in production.
  • Experience with observability platforms such as Datadog, and a strong intuition for what " good" monitoring looks like.
  • Exposure to Infrastructure as Code (Terraform) and GitOps-based deployment workflows (ArgoCD, GitHub Actions, or similar).
  • Familiarity with the broader technology stack: Java, Go, Kafka, Redis, Snowflake, and Postgres.
  • Strong analytical and problem-solving skills-you thrive on ambiguous, high-stakes production problems.
  • A product mindset applied to operational tooling: you think about usability, adoption, and documentation when building internal solutions.
  • Excellent communication skills and the ability to work fluidly across engineering, operations, and b
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Production Engineer
Senior Production Engineer

Clear Street • Greater London

On-site
GBP 90,000 - 130,000
Senior Production Engineer, Reliability & Automation
Senior Production Engineer, Reliability & Automation

Clear Street • Greater London

On-site
GBP 90,000 - 130,000
Production Engineer - Trading Systems
Production Engineer - Trading Systems

Orbis Group • Greater London

On-site
GBP 110,000 - 150,000
Senior Production Engineer: Reliability & Automation Lead
Senior Production Engineer: Reliability & Automation Lead

clearstreet • Greater London

On-site
GBP 110,000 - 140,000
Senior Platform Engineer
Senior Platform Engineer

Understanding Recruitment • Greater London

On-site
GBP 90,000 - 120,000
Lucrative Performance-based bonus
Equity package with significant long‑m
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CloudFactory • Reading

On-site
GBP 70,000 - 110,000
Platform Engineer (Mid-level)
Platform Engineer (Mid-level)

Team17 • Wakefield

On-site
GBP 55,000 - 75,000
Full Stack Engineer, Platform Reliability
Full Stack Engineer, Platform Reliability

Worky • East Midlands

On-site
GBP 65,000 - 95,000
Senior Solutions Engineer
Senior Solutions Engineer

Kroll • United Kingdom

On-site
GBP 80,000 - 110,000
Infrastructure Engineer
Infrastructure Engineer

Jobtailor • Greater London

Hybrid
GBP 90,000 - 130,000