Senior Sre Engineer - Dempo

Dempo

Málaga

Híbrido

EUR 60.000 - 90.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura completa en un minuto — currículum adaptado y carta de presentación, listos para enviar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Remote-friendly
Private health insurance
Learning budget
Learning time
Office Granada
Budget for certifications

Descripción de la vacante

Dempo is seeking a Senior SRE Engineer to lead production resilience within its infrastructure team in Málaga, Spain. You will own incident response, reliability metrics, and the observability strategy while mentoring cloud/platform engineers.

You will design SLIs/SLOs, drive CI/CD improvements, and balance speed with long-term reliability in a remote-first, office-flexible environment that offers private health insurance, learning funds, and certifications.

Formación

  • 5+ years of SRE/Platform engineering experience.
  • Strong understanding of incident response practices.
  • Deep knowledge of the observability stack: logs, metrics, traces.

Responsabilidades

  • Lead incidents as commander and manage severity.
  • Own incident records and link deployments to incidents.
  • Design and implement observability across the stack; define SLIs/SLOs.

Conocimientos

SRE experience
Kubernetes
Terraform
Bash/Python
Git
Incident response
Observability

Herramientas

Kubernetes
Terraform
CI/CD

Descripción del empleo

We are looking for a Senior SRE Engineer to join our infrastructure team and take technical leadership over production resilience. This role sits at the Senior level on our Cloud/Platform/SRE career path — reliability engineering with a heavy focus on metrics and production systems. You’ll define SLIs and SLOs, lead incident response as commander, drive observability strategy end to end, and mentor cloud/platform engineers as you go.

You’ll work closely with Product and Engineering, balancing speed, quality, and long-term reliability, while making the architectural calls that keep our systems resilient under load.

Responsibilities
Reliability & Incident Management
  • Lead incidents as commander: set and revise severity, and know when to mitigate first and diagnose later
  • Own the incident record and timeline standard, including the link between deployments and incidents
  • Communicate with stakeholders while an incident is active
  • Conduct blameless postmortems and drive toil identification and elimination as measured work
Observability
  • Implement the three pillars of observability (logs, metrics, traces) end to end
  • Design metrics and query strategy — recording rules, dashboard design that separates on-call needs from analyst needs
  • Define SLIs and SLOs for critical services, choosing the indicator that reflects user experience over the one that's easiest to measure
  • Design alerting systems — routing, escalation, deduplication, and alert fatigue reduction (multi-window burn-rate alerts)
Platform & Production Systems
  • Design workload health signals — liveness, readiness, and startup probes — and reason about workload lifecycle (SIGTERM handling, termination grace periods, connection draining)
  • Build runbook automation and self-healing systems to reduce operational toil
  • Contribute to CI/CD framework improvements and cost optimization initiatives
Technical Leadership
  • Make architectural decisions for reliability-critical systems
  • Mentor cloud/platform engineers
  • Influence technical direction on infrastructure and platform decisions
Requirements
  • 5+ years of experience in SRE, Cloud, or Platform engineering roles
  • Profound knowledge of SRE principles and incident response practices
  • Profound knowledge of the observability stack: log aggregation and query design, metrics/dashboard design, and distributed tracing
  • Profound knowledge of Kubernetes cluster operations and workload objects (Deployments, StatefulSets, Jobs, DaemonSets) and their failure modes
  • Solid to profound knowledge of networking fundamentals: DNS as infrastructure, TLS certificate lifecycle, load balancing and reverse proxies
  • Hands‑on experience with Infrastructure as Code (Terraform or equivalent)
  • Profound knowledge of process and OS‑level architecture trade‑offs as they apply to containers (immutable infrastructure, image strategy)
  • Scripting proficiency (Bash or Python) for tooling and automation
  • Strong Git and collaborative workflow experience
Nice to Have
  • Experience with chaos engineering or failure injection programs
  • Exposure to multi‑region or multi‑cloud design trade‑offs
  • Familiarity with service mesh implementations
  • Prior mentoring or technical leadership experience
  • Certifications such as Site Reliability Engineering (SRE) Foundation or Observability Foundation
  • Contributions to open‑source observability or Kubernetes tooling
What We Offer
  • Permanent contract.
  • Flexible working hours (core hours: 09:30 – 13:30).
  • Remote‑first culture, with the option to work from our Granada office.
  • 30 working days of annual leave, plus December 24th and December 31st as additional company days off that do not count against your holiday allowance.
  • Private health insurance.
  • Your choice of MacBook or Lenovo.
  • Unlimited access to learning platforms.
  • Learning time during working hours.
  • Budget for certifications and specialised training.
  • Employee referral programme.
  • Opportunity‑based bonuses.
  • Stable, long‑term projects.
  • Real opportunities for professional growth.
  • The opportunity to play a key role in the growth of a modern Software Engineering company.
Salary Range

Dempo is where technology, teamwork, and your professional growth come together.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior SRE Engineer
Senior SRE Engineer

Dempo • España

Presencial
PHP 6.531.000 - 10.160.000
Remote-first culture
Granada office option
Private health insurance
+4
Senior Sre Engineer - Tarragona
Senior Sre Engineer - Tarragona

Dempo • Tarragona

Híbrido
EUR 90.000 - 150.000
Private health insurance
Remote-friendly culture
Learning budget
+4
Remote-First Senior SRE Engineer - Lead Reliability
Remote-First Senior SRE Engineer - Lead Reliability

Dempo • Tarragona

Híbrido
EUR 90.000 - 150.000
Private health insurance
Remote-friendly culture
Learning budget
+4
SRE/DevOps Engineer in Barcelona
SRE/DevOps Engineer in Barcelona

Blu Selection • Barcelona

Híbrido
EUR 60.000 - 75.000
Hybrid work model
Senior SRE Leader — Observability & Reliability (Remote)
Senior SRE Leader — Observability & Reliability (Remote)

Dempo • Málaga

Híbrido
EUR 60.000 - 90.000
Remote-friendly
Private health insurance
Learning budget
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Comply365 • Barcelona

Presencial
EUR 100.000 - 120.000
Laptop & monitor
Annual learning budget
Conference budget
+2
Junior Cloud Engineer
Junior Cloud Engineer

Dempo • Granada

Híbrido
EUR 28.000 - 36.000
Private health insurance
Remote-first culture
Granada office option
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

vistairhr • Barcelona

Híbrido
EUR 100.000 - 120.000
Equipment: Laptop
Annual learning budget
Conference budget
+2
sre engineer for high-availability services
sre engineer for high-availability services

Enfint • Barcelona

Presencial
EUR 80.000 - 120.000
Meal compensation
Unlimited health insurance
Gym membership support
+3
Junior Aws Cloud Engineer
Junior Aws Cloud Engineer

Dempo • Arbo

Híbrido
EUR 25.000 - 35.000
Permanent contract
Flexible hours
Remote work (Granada office)
+9