Turn this role into an interview — a resume and cover letter built around what this employer wants.
Understanding Recruitment is seeking a Senior Site Reliability Engineer to enhance reliability, observability, and tooling for a latency-sensitive production platform. The role spans production infrastructure, monitoring, incident response, and deployment workflows with a strong Linux and networking focus.
You will work to improve CI/CD, automation, and internal tooling while elevating developer experience from local setups to production.
We're partnered with a technology company building high-performance infrastructure for decentralised financial markets.
They're looking for a Senior Site Reliability Engineer to improve the reliability, observability and operational tooling behind a latency-sensitive production platform.
The role covers production infrastructure, monitoring and alerting, incident diagnosis, deployment workflows, infrastructure automation and developer tooling. There is also a strong Linux and systems element, particularly around networking, host performance and running high-performance services in production.
The platform is still relatively early, so there is plenty of scope to improve how things are operated, introduce better automation and help set the standards the wider engineering team works to.
The systems are latency-sensitive, so the role can extend into areas such as host-level tuning, kernel settings, CPU isolation and networking behaviour.