Staff Site Reliability Engineer

Doghouse Recruitment

España

Remote

EUR 110,000 - 170,000

Full time

7 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Doghouse Recruitment is seeking a Senior/Staff Site Reliability Engineer to own reliability end-to-end in a bare-metal Linux/Data Center environment. This remote EU role focuses on reducing incidents, improving latency, and building automation to kill toil while maintaining deployment safety.

You will work near the metal across Linux, Kubernetes internals, and networking, with on-call duties as part of the role.

Qualifications

  • Extensive and recent production engineering experience on bare metal / on-prem / data center infra.
  • Deep hands-on Linux performance and debugging at kernel level.
  • Strong networking understanding (DNS/TCP/TLS, latency, loss, congestion).
  • Experience with Terraform, Docker, Helm and modern CI/CD practices.
  • Strong Go and/or Python coding skills beyond automation scripting.
  • Experience in low latency environments.

Responsibilities

  • Own production reliability end-to-end and define SLIs/SLOs.
  • Run error budget conversations and ship changes that reduce incidents.
  • Build automation to kill toil and improve deployment safety (canary/rollback).
  • Work near the metal across Linux, Kubernetes internals and networking.
  • Participate in on-call rotations and reduce on-call load.

Skills

Bare metal
Linux kernel
Networking
Kubernetes internals
Terraform
Docker
Helm
CI/CD
Go
Python
Low latency
On-call reduction

Tools

Terraform
Docker
Helm
Kubernetes

Job description

Staff Site Reliability Engineer – Bare Metal Linux – Data Center – Networking

Location: 100% remote within the EU

Our client is building a cloud platform for high-throughput, compute-heavy workloads. They operate large-scale infrastructure where failure modes are real, capacity is finite, and reliability needs to be engineered, not “handled”.

We’re seeking a Senior/Staff SRE who will own production reliability end-to-end for our client: define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency (p95/p99). You’ll build automation to kill toil, improve deployment safety (canary/rollback), and turn observability into signal rather than noise.

This is a bare-metal environment: think Linux, datacenters, physical fleets, and real hardware constraints, not managed services. You’ll work close to the metal across Kubernetes internals (scheduling, autoscaling behavior, kubelet pressure/evictions, etcd/control plane), Linux performance (CPU/memory/I/O contention), and network debugging (DNS/TCP/TLS, packet loss, congestion). On-call is part of the job, but success is measured by how much you reduce it.

Must requirements:

  • Extensive and recent Production Engineering experience running bare metal / on-prem / data center infrastructure (not public cloud only)
  • Deep hands-on expertise in Linux systems debugging and performance at a kernel level(CPU, memory, I/O, low-level behaviors)
  • Strong understanding of networking (DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting under load)
  • Experience with Terraform, Docker, Helm, and modern CI/CD practices
  • Strong coding skills are required for this role either in Go, and/or Python, beyond automation scripting - Real engineering capability is a must
  • Experience in Low Latency environments.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff SRE – Remote EU (Bare-Metal Linux & DC Networking)
Staff SRE – Remote EU (Bare-Metal Linux & DC Networking)

Doghouse Recruitment • Spain

Remote
EUR 110,000 - 170,000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Madrid

On-site
EUR 45,000 - 60,000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Ribarroja del Turia

On-site
EUR 40,000 - 70,000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Exciting projects: Modern solutions with Fortune 500 and top product companies
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hydrolix • Spain

On-site
EUR 70,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

Emburse, Inc. • Barcelona

On-site
EUR 90,000 - 130,000
Flexible spending accounts
Generous paid time off
Paid parental leave
+9
Senior Middle Level DevOps Engineer Full time - Hybrid Remote
Senior Middle Level DevOps Engineer Full time - Hybrid Remote

kentech-sp • Spain

Hybrid
EUR 55,000 - 80,000
Gym membership
Private health insurance
Pension plan
Senior SRE
Senior SRE

Morgan Philips Specialist Recruitment • Madrid

On-site
EUR 80,000 - 120,000
Hybrid work in Madrid
Permanent contract
Flexible working hours
+1
Lead Site Reliability Engineer - Remote and Flexible
Lead Site Reliability Engineer - Remote and Flexible

SITA Group • Barcelona

Hybrid
EUR 90,000 - 130,000
Flex Week (WFH 2 days)
Flex Day
Flex Location (30 days remote)
+3
Senior DevOps Engineer
Senior DevOps Engineer

Jobgether SRL • Spain

Remote
EUR 70,000 - 120,000
Fully remote Europe
International teams
Technical ownership
+6
Senior SRE Engineer
Senior SRE Engineer

Codeway • Barcelona

On-site
EUR 90,000 - 120,000