This role owns performance and reliability foundations: infrastructure, runtime systems, monitoring, and tooling that keeps everything stable and observable. You’ll build for predictable behavior under stress, with disciplined change management and tight feedback loops.
Outcomes (First 90 Days)
- Improve a measurable reliability/performance bottleneck (latency/jitter/stability/recovery).
- Implement or upgrade monitoring for a critical path.
- Improve an operational workflow/runbook to reduce recurring firefights.
Responsibilities
- Engineer and maintain core infrastructure: compute, networking, deployment paths, monitoring.
- Improve latency/jitter and stability across critical components.
- Build tooling for visibility: metrics, dashboards, alert routing, incident workflows.
- Coordinate constraints and priorities with dev/research/trading.
- Maintain operational readiness: runbooks, escalation paths, failure‑mode thinking.
Required Skills (Must‑Have)
- Experience building reliable production systems with monitoring and incident handling.
- Ability to reason about trade‑offs under constraints.
Bonus (Nice‑To‑Have)
- Low‑latency work (profiling, efficient I/O, tuning).
- Experience with colocated environments and redundancy patterns.
- Strong coding ability (Python and/or C++/Rust).
How We Work
Flat teams, clear ownership, evidence‑based reviews.
Problems are solved systematically, not politically.
How To Apply
Apply using the form or email us. We read every application. If there's a fit, you'll hear back from us.