We are looking for a Senior SRE Engineer to join our infrastructure team and take technical leadership over production resilience. This role sits at the Senior level on our Cloud/Platform/SRE career path — reliability engineering with a heavy focus on metrics and production systems. You’ll define SLIs and SLOs, lead incident response as commander, drive observability strategy end to end, and mentor cloud/platform engineers as you go.
You’ll work closely with Product and Engineering, balancing speed, quality, and long-term reliability, while making the architectural calls that keep our systems resilient under load.
Responsibilities
Reliability & Incident Management
- Lead incidents as commander: set and revise severity, and know when to mitigate first and diagnose later
- Own the incident record and timeline standard, including the link between deployments and incidents
- Communicate with stakeholders while an incident is active
- Conduct blameless postmortems and drive toil identification and elimination as measured work
Observability
- Implement the three pillars of observability (logs, metrics, traces) end to end
- Design metrics and query strategy — recording rules, dashboard design that separates on-call needs from analyst needs
- Define SLIs and SLOs for critical services, choosing the indicator that reflects user experience over the one that's easiest to measure
- Design alerting systems — routing, escalation, deduplication, and alert fatigue reduction (multi-window burn-rate alerts)
Platform & Production Systems
- Design workload health signals — liveness, readiness, and startup probes — and reason about workload lifecycle (SIGTERM handling, termination grace periods, connection draining)
- Build runbook automation and self-healing systems to reduce operational toil
- Contribute to CI/CD framework improvements and cost optimization initiatives
Technical Leadership
- Make architectural decisions for reliability-critical systems
- Mentor cloud/platform engineers
- Influence technical direction on infrastructure and platform decisions
Requirements
- 5+ years of experience in SRE, Cloud, or Platform engineering roles
- Profound knowledge of SRE principles and incident response practices
- Profound knowledge of the observability stack: log aggregation and query design, metrics/dashboard design, and distributed tracing
- Profound knowledge of Kubernetes cluster operations and workload objects (Deployments, StatefulSets, Jobs, DaemonSets) and their failure modes
- Solid to profound knowledge of networking fundamentals: DNS as infrastructure, TLS certificate lifecycle, load balancing and reverse proxies
- Hands-on experience with Infrastructure as Code (Terraform or equivalent)
- Profound knowledge of process and OS-level architecture trade-offs as they apply to containers (immutable infrastructure, image strategy)
- Scripting proficiency (Bash or Python) for tooling and automation
- Strong Git and collaborative workflow experience
Nice to Have
- Experience with chaos engineering or failure injection programs
- Exposure to multi-region or multi-cloud design trade-offs
- Familiarity with service mesh implementations
- Prior mentoring or technical leadership experience
- Certifications such as Site Reliability Engineering (SRE) Foundation or Observability Foundation
- Contributions to open-source observability or Kubernetes tooling
What We Offer
- Permanent contract.
- Flexible working hours (core hours: 09:30 – 13:30).
- Remote-first culture, with the option to work from our Granada office.
- 30 working days of annual leave, plus December 24th and December 31st as additional company days off that do not count against your holiday allowance.
- Private health insurance.
- Your choice of MacBook or Lenovo.
- Unlimited access to learning platforms.
- Learning time during working hours.
- Budget for certifications and specialised training.
- Employee referral programme.
- Opportunity-based bonuses.
- Stable, long-term projects.
- Real opportunities for professional growth.
- The opportunity to play a key role in the growth of a modern Software Engineering company.
Selection Process
- With pleasure, we receive your CV and we give you a call
- Now it’s when we put faces to names, we’d love to get to chat with you!
- Let’s deepen a bit more with a technical interview, a chance to meet your People Partner
- We get back to you with offer and feed
Dempo is where technology, teamwork, and your professional growth come together.