Senior SRE Engineer

Dempo

España

On-site

PHP 6,531,000 - 10,160,000

Full time

45 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote-first culture
Granada office option
Private health insurance
Learning platforms access
Budget for certifications
Learning time during working hours
Core hours 09:30–13:30

Job summary

Dempo is seeking a Senior SRE Engineer to lead production resilience and drive reliability across our cloud/platform stack. You will own incident response, observability strategy, and architectural decisions that keep systems resilient under load.

You will mentor engineers, collaborate with product and engineering, and shape the reliability roadmap with a focus on measurable user experience and long-term stability.

Qualifications

  • 5+ years in SRE, Cloud, or Platform engineering roles.
  • Proficient in SRE principles and incident response practices.
  • Deep knowledge of the observability stack: logs, metrics, traces.

Responsibilities

  • Lead incidents as commander and manage incident timelines.
  • Define SLIs and SLOs for critical services and design alerting.
  • Design workload health signals and self-healing automation.
  • Mentor cloud/platform engineers and influence infrastructure decisions.

Skills

SRE fundamentals
Incident response
Observability
Kubernetes ops
Networking fundamentals
Terraform
Bash or Python
Git workflows

Tools

CI/CD pipelines

Job description

We are looking for a Senior SRE Engineer to join our infrastructure team and take technical leadership over production resilience. This role sits at the Senior level on our Cloud/Platform/SRE career path — reliability engineering with a heavy focus on metrics and production systems. You’ll define SLIs and SLOs, lead incident response as commander, drive observability strategy end to end, and mentor cloud/platform engineers as you go.

You’ll work closely with Product and Engineering, balancing speed, quality, and long-term reliability, while making the architectural calls that keep our systems resilient under load.

Responsibilities
Reliability & Incident Management
  • Lead incidents as commander: set and revise severity, and know when to mitigate first and diagnose later
  • Own the incident record and timeline standard, including the link between deployments and incidents
  • Communicate with stakeholders while an incident is active
  • Conduct blameless postmortems and drive toil identification and elimination as measured work
Observability
  • Implement the three pillars of observability (logs, metrics, traces) end to end
  • Design metrics and query strategy — recording rules, dashboard design that separates on-call needs from analyst needs
  • Define SLIs and SLOs for critical services, choosing the indicator that reflects user experience over the one that's easiest to measure
  • Design alerting systems — routing, escalation, deduplication, and alert fatigue reduction (multi-window burn-rate alerts)
Platform & Production Systems
  • Design workload health signals — liveness, readiness, and startup probes — and reason about workload lifecycle (SIGTERM handling, termination grace periods, connection draining)
  • Build runbook automation and self-healing systems to reduce operational toil
  • Contribute to CI/CD framework improvements and cost optimization initiatives
Technical Leadership
  • Make architectural decisions for reliability-critical systems
  • Mentor cloud/platform engineers
  • Influence technical direction on infrastructure and platform decisions
Requirements
  • 5+ years of experience in SRE, Cloud, or Platform engineering roles
  • Profound knowledge of SRE principles and incident response practices
  • Profound knowledge of the observability stack: log aggregation and query design, metrics/dashboard design, and distributed tracing
  • Profound knowledge of Kubernetes cluster operations and workload objects (Deployments, StatefulSets, Jobs, DaemonSets) and their failure modes
  • Solid to profound knowledge of networking fundamentals: DNS as infrastructure, TLS certificate lifecycle, load balancing and reverse proxies
  • Hands-on experience with Infrastructure as Code (Terraform or equivalent)
  • Profound knowledge of process and OS-level architecture trade-offs as they apply to containers (immutable infrastructure, image strategy)
  • Scripting proficiency (Bash or Python) for tooling and automation
  • Strong Git and collaborative workflow experience
Nice to Have
  • Experience with chaos engineering or failure injection programs
  • Exposure to multi-region or multi-cloud design trade-offs
  • Familiarity with service mesh implementations
  • Prior mentoring or technical leadership experience
  • Certifications such as Site Reliability Engineering (SRE) Foundation or Observability Foundation
  • Contributions to open-source observability or Kubernetes tooling
What We Offer
  • Permanent contract.
  • Flexible working hours (core hours: 09:30 – 13:30).
  • Remote-first culture, with the option to work from our Granada office.
  • 30 working days of annual leave, plus December 24th and December 31st as additional company days off that do not count against your holiday allowance.
  • Private health insurance.
  • Your choice of MacBook or Lenovo.
  • Unlimited access to learning platforms.
  • Learning time during working hours.
  • Budget for certifications and specialised training.
  • Employee referral programme.
  • Opportunity-based bonuses.
  • Stable, long-term projects.
  • Real opportunities for professional growth.
  • The opportunity to play a key role in the growth of a modern Software Engineering company.
Selection Process
  1. With pleasure, we receive your CV and we give you a call
  2. Now it’s when we put faces to names, we’d love to get to chat with you!
  3. Let’s deepen a bit more with a technical interview, a chance to meet your People Partner
  4. We get back to you with offer and feed

Dempo is where technology, teamwork, and your professional growth come together.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Junior Cloud Engineer
Junior Cloud Engineer

Dempo • España

On-site
PHP 2,633,000 - 4,013,000
Permanent contract
Flexible hours
Remote-first culture
+7
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Senior Site Reliability Engineer (SRE) Operations
Senior Site Reliability Engineer (SRE) Operations

Teciem • Hinoba-an

Hybrid
PHP 2,772,000 - 4,291,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia • Philippines

On-site
PHP 1,000,000 - 1,800,000
Fixed compensation
Vacation leaves
Professional development opportunities
+3
Site Reliability Engineer (SRE) Operations
Site Reliability Engineer (SRE) Operations

Teciem • Hinoba-an

Hybrid
PHP 990,000 - 1,518,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Senior SRE Engineer: Lead Reliability & Observability (Remote)
Senior SRE Engineer: Lead Reliability & Observability (Remote)

Dempo • España

On-site
PHP 6,531,000 - 10,160,000
Remote-first culture
Granada office option
Private health insurance
+4
DevOps Engineer / Site Reliability Engineer (SRE)
DevOps Engineer / Site Reliability Engineer (SRE)

Concentrix • Mexico

On-site
PHP 1,000,000 - 1,600,000
Site Reliability Engineer
Site Reliability Engineer

EROAD Limited • Manila

On-site
PHP 900,000 - 1,500,000
Staff Site Reliability Engineer – Cloud Efficiency
Staff Site Reliability Engineer – Cloud Efficiency

Super • España

On-site
PHP 1,200,000 - 1,600,000
Medical / Health Insurance
Employee Assistance Programme