Stand out for this role — generate a tailored resume and cover letter in about a minute.
Clear Street is looking for a Production Engineer to own the reliability of our cloud-native platform. You will design monitoring, automate recovery, and partner with engineering and operations to reduce toil while increasing platform resilience.
You’ll balance on-call responsiveness with building scalable tooling and self-healing capabilities. You’ll work across teams to improve MTTR, implement robust observability with Datadog, and contribute to IaC and GitOps practices, ensuring safer
Clear Street’s mission is to give every sophisticated investor access to every asset, in every market, through a unified platform built for speed, transparency and scale.
We give our clients the technology, tools, and service once reserved for the largest institutions, rebuilt with modern infrastructure. Our single, cloud-native, end-to-end capital markets platform powers investor growth today and is transforming how they can interact with markets tomorrow.
As a Production Engineer, you sit at the intersection of software reliability and operational excellence. You own the health, resilience, and recovery of our production systems—while spending equal energy innovating solutions that eliminate human toil, reduce incident blast radius, and raise the reliability bar across the entire platform. You will partner closely with engineering, operations, and business teams to understand daily pain points and translate them into lasting automated solutions. Half your time is spent in the trenches—supporting production, responding to incidents, and deeply understanding how our systems behave under real conditions. The other half is yours to build: automation, tooling, and observability platforms that make tomorrow's on-call shift meaningfully easier than today's.
We believe resilient systems are built by engineers who understand them end to end. Our Production Engineering team is the first and last line of defense for our production platform. We treat reliability as a product, with uptime and engineer experience as our north stars. We combine the discipline of SRE with a builder's mindset: when we see a recurring problem, we build a solution—not a workaround.
You will work across every engineering and operations team to understand failure modes, quantify reliability gaps, and build platform capabilities that scale with the organization. Whether it's reducing MTTR from hours to minutes, building self-service diagnostic tools, or designing proactive alerting that catches issues before customers notice, your work will have immediate, measurable impact.
If you're passionate about making production systems invisible to end users—and you get energy from both firefighting and building the systems that make fires less likely—you'll thrive here.
We're looking for engineers who combine operational instinct with a builder's discipline.
You'll operate and build on a modern cloud-native platform that includes:
Within your first year, you'll have made a measurable impact on production reliability. Success l ooks like:
Your impact won't be measured by the number of incidents you respond to—it will be measured by how reliably our systems run and how quickly we recover when they don't.