Turn this role into an interview — a resume and cover letter built around what this employer wants.
Doghouse Recruitment is seeking a Senior/Staff Site Reliability Engineer to own reliability end-to-end in a bare-metal Linux/Data Center environment. This remote EU role focuses on reducing incidents, improving latency, and building automation to kill toil while maintaining deployment safety.
You will work near the metal across Linux, Kubernetes internals, and networking, with on-call duties as part of the role.
Location: 100% remote within the EU
Our client is building a cloud platform for high-throughput, compute-heavy workloads. They operate large-scale infrastructure where failure modes are real, capacity is finite, and reliability needs to be engineered, not “handled”.
We’re seeking a Senior/Staff SRE who will own production reliability end-to-end for our client: define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency (p95/p99). You’ll build automation to kill toil, improve deployment safety (canary/rollback), and turn observability into signal rather than noise.
This is a bare-metal environment: think Linux, datacenters, physical fleets, and real hardware constraints, not managed services. You’ll work close to the metal across Kubernetes internals (scheduling, autoscaling behavior, kubelet pressure/evictions, etcd/control plane), Linux performance (CPU/memory/I/O contention), and network debugging (DNS/TCP/TLS, packet loss, congestion). On-call is part of the job, but success is measured by how much you reduce it.
Must requirements: