Lead Site Reliability Engineer/ Expert

SITA

Barcelona

Hybrid

EUR 70,000 - 100,000

Full time

15 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Flex Week: Work from home up to 2 days
Flex-Location: up to 30 days remote
Employee wellbeing programs
Professional development opportunities
Competitive benefits

Job summary

SITA is seeking a Lead Site Reliability Engineer to drive proactive support and performance for critical products. You will identify root causes of incidents, implement automated workflows, and ensure smooth deployment with minimal disruption.

This role emphasizes automation, cross-functional collaboration, and ongoing reliability improvements across cloud and on-prem environments. Join a global team and help optimize incident response, problem management, and event handling while aligning with

Qualifications

  • Bachelor’s degree required in CS/IT/Engineering.
  • 6+ years in IT operations, service management, or infrastructure management.
  • Proven experience with high-availability systems and operational reliability.
  • Extensive RCA, incident management, and permanent solution development.
  • Hands-on with CI/CD, automation, monitoring, and IaC.

Responsibilities

  • Define, build, and maintain support systems to ensure high availability and performance.
  • Handle complex PSO cases and automate provisioning, self-healing, deployment, and monitoring.
  • Perform incident response and root cause analysis (RCA) for critical system failures.
  • Monitor system performance and establish SLIs/SLOs; implement reliability best practices.
  • Coordinate with Product T&E, ICE, and Service Architects for productization as SGS technical expert.
  • Ensure Operations readiness to support new products.
  • Accountable for in-scope product availability and performance.

Skills

RCA & incident management
CI/CD pipelines
Automation
SRE mindset
Cross-functional collaboration

Education

Bachelor’s degree in Computer Science / IT / Engineering

Tools

AKS
Kubernetes
Ansible
Python / Bash scripting
Terraform
Azure / AWS

Job description

Overview

WELCOME TO SITA At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry. You'll find us in 95% of international airports, working closely with over 2,500 transportation and government clients. Each partnership brings unique challenges, and we thrive on delivering fresh solutions and cutting-edge tech to keep operations running like clockwork. We don't just move the world forward-we're proud to be recognized as a Great Place to Work® by 79% of our employees and certified in most of our growing locations. Here, we feel empowered, supported, and inspired to grow. Are you ready to love your job? The adventure begins right here, with you, at SITA.

About The Role And The Team

As Lead Site Reliability Engineer/ Expert you will be responsible for the proactive support of products so that there is high product performance that is continuously improved. Responsible for identifying and resolving the root causes of operational incidents, implementing solutions to improve stability and prevent recurrence. Manages the creation and maintenance of the event catalog to trigger events and develop both manual remediation approaches and automated workflows to resolve alerts. Oversee the deployment of IT services and solutions ensuring successful integration with minimal disruption. Focuses on operational automation and integration to enhance efficiency and collaboration between development and operations within service operations.

What You Will Do
  • Define, build, and maintain support systems to ensure high availability and performance.
  • Handle complex cases for the PSO.
  • Implement automation for system provisioning, self-healing, auto-recovery, deployment, and monitoring.
  • Perform incident response and root cause analysis (RCA) for critical system failures.
  • Monitor system performance and establish Service-Level Indicators (SLIs) and Service-Level Objectives (SLOs).
  • Collaborate with Development and Operations to integrate reliability best practices, including zero-downtime architecture.
  • Proactively identify and remediate performance issues.
  • Work closely with Product T&E, ICE, and Service Architects for new product productization as SGS technical expert.
  • Coordinate with internal and external stakeholders to improve service performance and ensure high availability.
  • Ensure Operations readiness to support new products.
  • Accountable within SGS for in-scope product availability and performance.
Problem Management
  • Conduct thorough problem investigations and root cause analyses to diagnose recurring incidents and service disruptions.
  • Coordinate with Incident Management teams and collaborate with PSOs and Engineering/Product teams to implement permanent solutions.
  • Monitor effectiveness of problem resolution activities and provide regular reporting to ensure continuous improvement.
Event Management
  • Define, build, and maintain an event catalog specifying active events, thresholds, and remediation actions; optimize it for efficiency.
  • Develop event response protocols, provide training, and ensure efficient incident handling.
Customer & Operational Support
  • Collaborate with Customer Success Managers to implement initiatives that enhance customer satisfaction and retention.
  • Prepare reports, documentation, and communication materials covering customer metrics, updates, and product changes.
  • Identify and implement improvements in internal processes and workflows.
  • Contribute to knowledge management resources such as FAQs and training materials.
Data Steward Responsibilities
  • Implement data governance policies defined by the Data Owner and ensure adherence to standards.
  • Monitor data quality, consistency, and compliance on an ongoing basis.
  • Act as a Subject Matter Expert (SME) for data within the assigned area, providing guidance and answering queries.
Qualifications
About Your Skills
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
  • 6+ years of experience in IT operations, service management, or infrastructure management, including roles such as Site Reliability Engineer, Problem Manager, or DevOps Manager.
  • Proven experience managing high-availability systems and ensuring operational reliability.
  • Extensive experience in root cause analysis (RCA), incident management, and developing permanent solutions for recurring service disruptions.
  • Hands‑on experience with CI/CD pipelines, automation, system performance monitoring, and infrastructure as code (IaC).
  • Strong background in collaborating with cross-functional teams (Development, Operations, Engineering, etc.) to improve operational processes and service delivery.
  • Experience managing deployments, conducting risk assessments, and optimizing event and problem management processes.
  • Familiarity with cloud technologies, containerization, and scalable architectures, including zero‑downtime deployment strategies.
Technical Skills (Must-to-Have)
  • Strong AKS & On prem K8s skills and experience,
  • Scripting (Ansible & Bash, Python - combination of anything would be great),
  • Automation,
  • CI/CD pipeline,
  • Terraform exposure,
  • Azure (or) AWS skill.
  • Basic DB skills.
  • Strong problem-solving skills & quick learner.
  • SRE mindset.
Please note:

Person will be working with global team, so he /she has to be flexible for overlap or stretch for operational urgency.

What We Offer

We're all about diversity. We operate in 200 countries and speak 60 different languages and cultures. We're really proud of our inclusive environment. Our offices are comfortable and fun places to work, and we make sure you get to work from home too. Find out what it's like to join our team and take a step closer to your best life ever.

Flex Week: Work from home up to 2 days/week (depending on your team's needs)

Flex Day: Make your workday suit your life and plans.

Flex-Location: Take up to 30 days a year to work from any location in the world.

Employee Wellbeing: We have got you covered with our Employee Assistance Program (EAP), for you and your dependents 24/7, 365 days/year. We also offer Champion Health - a personalized platform that supports a range of wellbeing needs.

Professional Development: At SITA, we believe growth fuels innovation. Our learning ecosystem offers access to world‑class platforms and programs designed to help you thrive. From LinkedIn Learning, Microsoft's Enterprise Skills Initiative, and Airport Council International -available to all employees‑to specialized solutions like Pluralsight for technology upskilling, Harvard Business Publishing for people leadership, Stanford for strategic development and many others, we align learning opportunities with your Development Plan and our business priorities. Your development journey is supported every step of the way.

Competitive Benefits: Competitive benefits that make sense with both your local market and employment status.

SITA is an Equal Opportunity Employer. We value a diverse workforce. In support of our Employment Equity Program, we encourage women, aboriginal people, members of visible minorities, and/or persons with disabilities to apply and self-identify in the application process.

Salary / Compensation Note

Hidden (-999)

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA Group • Barcelona

Hybrid
EUR 90,000 - 130,000
Flex Week (WFH 2 days)
Flex Day
Flex Location (30 days remote)
+3
Lead Technical Analyst
Lead Technical Analyst

SITA Group • Barcelona

On-site
EUR 90,000 - 120,000
Flex Week
Flex Day
Flex-Location
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

SITA • Barcelona

On-site
EUR 60,000 - 80,000
Flex Week: Work from home up to 2 days/week
Flex Day: Make your workday suit your life
Flex-Location: Work from any location for 30 days a year
Pre-Sales Senior Manager
Pre-Sales Senior Manager

SITA • Barcelona

Hybrid
EUR 90,000 - 120,000
Flex Week
Flex Day
Flex-Location
+3
Senior Software Developer (.NET)
Senior Software Developer (.NET)

SITA Group • Barcelona

On-site
EUR 90,000 - 120,000
Flex Week
Flex Day
Flex-Location
+3
Senior Data Engineer/Expert/ Specialist
Senior Data Engineer/Expert/ Specialist

SITA Group • Barcelona

On-site
EUR 90,000 - 130,000
Flex Week
Flex Day
Flex Location
+3
Customer Success Senior Consultant
Customer Success Senior Consultant

SITA • Barcelona

On-site
EUR 48,000 - 72,000
Senior Business Consultant
Senior Business Consultant

SITA • Madrid

On-site
EUR 60,000 - 85,000
Work from home options
Flexible work hours
Employee Wellbeing Program
+2
Senior Analyst Service Level & Supplier Mgmt
Senior Analyst Service Level & Supplier Mgmt

SITA • Barcelona

Hybrid
EUR 52,000 - 76,000
Flex Week
Flex Day
Flex-Location
+3
Project Manager
Project Manager

SITA Group • Barcelona

Hybrid
EUR 70,000 - 110,000
Flex Week
Flex Day
Flex-Location
+3