Resilience Engineer: Chaos & Reliability Architect

Vodafone

Lisboa

Hybrid

EUR 60,000 - 85,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Phone plan discounts
Well-being program
Learning resources

Job summary

Vodafone in Portugal is seeking a Senior Site Reliability Engineer to design and govern resilience strategies across system architecture, deployment, monitoring and incident response. You will define SLOs/SLIs, track error budgets, and collaborate with performance and operations teams to meet targets.

In this role you will implement fault injection, chaos engineering, and scenario simulations to validate platform robustness, while partnering with product, infra, architecture and development

Qualifications

  • BSc degree in Software Engineering, Computer Science or a related technical discipline, or equivalent professional experience.
  • Strong expertise in Site Reliability Engineering principles, including SLOs/SLIs, error budgets, observability and systemic reliability improvement
  • Deep understanding of distributed-system behaviour, failure modes and dependency management
  • Strong experience analysing system performance, including latency, throughput, errors, saturation and service degradation
  • Experience designing resilience strategies for highly available and fault-tolerant systems
  • Experience defining and validating failover, degraded-mode and recovery scenarios
  • Strong knowledge of telemetry and observability platforms such as Grafana, Prometheus, OpenTelemetry, Loki, Splunk or equivalent
  • Experience with Synthetic Monitoring and service-level performance monitoring; RUM experience is advantageous where relevant
  • Experience applying AI-assisted observability and automation capabilities to support anomaly detection, incident analysis, event correlation, root-cause investigation and predictive reliability, while maintaining appropriate human oversight and evidence-based decision-making
  • Proven ability to analyse incidents, identify systemic failure patterns and drive cross-team improvement
  • Experience with resilience testing, failure injection or Chaos Engineering is desirable
  • Understanding of Business Continuity and Disaster Recovery principles and standards such as ISO 22301 is desirable
  • Ability to translate complex technical findings into clear business and executive recommendations
  • Strong analytical and systems-thinking capability
  • Ability to challenge assumptions using evidence and data
  • Strong stakeholder management and influencing skills
  • Clear presentation and communication skills for technical, business and senior-management audiences
  • Ability to work across organisational boundaries without relying on direct authority
  • Strong ownership, autonomy, prioritisation and time-management skills
  • Fluent in written and spoken English

Responsibilities

  • Developing and governing resilience strategies across system architecture, deployment, monitoring, and incident response
  • Defining and tracking stability KPIs (e.g., MTTD, MTTR, error budgets), partnering with performance and operations teams to meet or exceed targets
  • Designing and implementing fault injection testing, chaos engineering practices, and scenario-based simulations to validate platform robustness
  • Collaborating with product, infrastructure, architecture and development teams to re-design services with built-in redundancy, failover, and graceful degradation
  • Driving automation and observability improvements to reduce noise, increase fault detection speed, and support predictive failure mitigation
  • Contributing to the design and maintenance of our Business Continuity and Disaster Recovery Plan (BCDR), ensuring IoT systems remain resilient and recoverable in the face of unexpected disruptions.
  • Owning the resilience roadmap and continuously assessing emerging threats, technologies, and architectural shifts to guide evolution of stability practices
  • Evangelizing a culture of resilience through internal communication, workshops, and post-incident learning programs

Skills

SRE principles
SLOs/SLIs
Observability
Grafana/Prometheus/OpenTelemetry/Loki
Incident response
English proficiency
Chaos engineering
Telemetry platforms

Education

BSc in Software Eng/CS or related

Tools

Grafana
Prometheus
OpenTelemetry
Loki
Splunk

Job description

Vodafone in Portugal is seeking a Senior Site Reliability Engineer to design and govern resilience strategies across system architecture, deployment, monitoring and incident response. You will define SLOs/SLIs, track error budgets, and collaborate with performance and operations teams to meet targets.

In this role you will implement fault injection, chaos engineering, and scenario simulations to validate platform robustness, while partnering with product, infra, architecture and development

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Resilience Engineer: Chaos Testing & Observability Expert
Resilience Engineer: Chaos Testing & Observability Expert

Vodafone Group Plc • Lisboa

On-site
EUR 85,000 - 120,000
Hybrid Work Model
Mobile plan discounts
Learning and development
+2
Resilience Engineer
Resilience Engineer

Vodafone • Lisboa

Hybrid
EUR 60,000 - 85,000
Hybrid work model
Phone plan discounts
Well-being program
+1
Resilience-Focused SRE for Telco Systems
Resilience-Focused SRE for Telco Systems

Decskill • Lisboa

Hybrid
EUR 60,000 - 88,000
Resilience Engineer
Resilience Engineer

Vodafone Group Plc • Lisboa

On-site
EUR 85,000 - 120,000
Hybrid Work Model
Mobile plan discounts
Learning and development
+2
Remote Site Reliability Engineer: Resilience & Observability
Remote Site Reliability Engineer: Resilience & Observability

Intermedia Intelligent Communications • Portugal

On-site
EUR 40,000 - 70,000
Cloud & Telco Infrastructure Engineer (OpenShift & VMware)
Cloud & Telco Infrastructure Engineer (OpenShift & VMware)

Vodafone • Lisboa

On-site
EUR 60,000 - 90,000
Senior Home Device Architect & Engineer (Hybrid)
Senior Home Device Architect & Engineer (Hybrid)

Vodafone • Portugal

Hybrid
EUR 70,000 - 100,000
Hybrid Work Model
Mobile phone & data plan
Discounts on services and products
+4
Network Detection & Response Solution Engineer
Network Detection & Response Solution Engineer

Vodafone Group Plc • Amadora

Hybrid
EUR 60,000 - 90,000
Hybrid work model
Employee discounts
Learning resources
+1
E2E OSS Solution Architect
E2E OSS Solution Architect

Vodafone Group Plc • Lisboa

Hybrid
EUR 60,000 - 90,000
Hybrid Work Model
Vodafone Products and Services
Recognition
+3
Hybrid IP Network Engineer - Connectivity & Automation
Hybrid IP Network Engineer - Connectivity & Automation

Vodafone • Almada

Hybrid
EUR 50,000 - 80,000
Hybrid Work Model
Vodafone Products and Services
Recognition programs
+3