Lead Site Reliability Engineer

ISA

Dubai

On-site

AED 350,000 - 700,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ISA seeks a Senior Site Reliability Engineer in Dubai to lead reliability for Java-based applications and airline reservation systems, implementing SRE best practices and proactive monitoring. You will own incident management, CI/CD, and capacity planning while mentoring teams for operational excellence.

You will collaborate with software and infra teams to ensure production readiness, drive performance improvements, and maintain high availability across mission-critical services.

Qualifications

  • 6–8 years of experience in Software Engineering, SRE, or Production Support for Java applications.

Responsibilities

  • Lead SRE initiatives for Java-based microservices and monolithic applications.
  • Define and manage SLAs, SLOs, and error budgets.
  • Lead RCA for production incidents and implement preventive measures.
  • Drive architectural improvements for performance, scalability, and high availability.
  • Mentor teams on incident management and production support.
  • Collaborate with engineering and infra teams to ensure production readiness.
  • Define CI/CD standards and GitOps practices across teams.
  • Set Docker/Kubernetes standards for scalability and resilience.
  • Own on-call escalation framework and post-incident reviews.
  • Evaluate new tools to strengthen reliability and capacity planning.
  • Report reliability metrics to management.

Skills

Java
Spring Boot
Troubleshooting
MS Office
SQL
Distributed systems
Monitoring/Observability
Prometheus
Grafana
Elasticsearch
Datadog

Education

Bachelor’s Degree in Computer Engineering/ Computer Science/ Information Technology

Tools

Prometheus
Grafana
Elasticsearch
Datadog
Docker
Kubernetes
JBoss/Application Server

Job description

Job Description:

Job Purpose

To lead the reliability, availability, performance, and continuous improvement of Air Arabia's mission-critical Java-based applications and airline reservation systems by implementing Site Reliability Engineering (SRE) best practices. Responsible for driving operational excellence through proactive monitoring, automation, incident management, scalability improvements, and collaboration with cross-functional teams, while ensuring compliance with organizational policies, industry standards, and applicable regulatory requirements.

Key Result Responsibilities
  • Lead Site Reliability Engineering (SRE) initiatives for Java-based microservices and monolithic applications supporting mission-critical airline operations and Passenger Service Systems (PSS).
  • Establish, monitor, and continuously improve system reliability by defining and managing Service Level Agreements (SLAs), Service Level Objectives (SLOs), and error budgets.
  • Lead and facilitate root cause analysis (RCA) for complex production incidents, ensuring timely resolution and implementation of preventive measures to minimize recurrence.
  • Drive architectural enhancements to improve system performance, scalability, resilience, availability, and operational efficiency through optimization techniques such as caching and distributed system design.
  • Mentor and provide technical guidance to team members by promoting best practices in incident management, production support, troubleshooting, and operational excellence.
  • Collaborate closely with software engineering, infrastructure, and cross-functional teams to enhance application design, improve system reliability, and ensure production readiness.
  • Own and define CI/CD pipeline standards, GitOps practices, and infrastructure automation strategy, setting the framework that engineers across the team execute against.
  • Set the containerization and orchestration strategy (Docker, Kubernetes) for the team, defining standards for scalability, resilience, and high availability that other engineers implement.
  • Own the on‑call escalation framework, ensuring adequate coverage and clear escalation paths, and lead post‑incident reviews.
  • Evaluate new tools and technologies to strengthen reliability and represent SRE in capacity planning and release readiness reviews.
  • Own reliability reporting to management, translating SLA/SLO performance and incident trends into clear business updates.
Qualifications (Academic, Training, Languages)
  • Bachelor’s Degree in Computer Engineering/ Computer Science/ Information Technology.
  • Fluent in English Language.
  • Strong expertise in Java, Spring Boot, and troubleshooting complex production issues within enterprise environments.
  • Airline, aviation, or travel industry experience, particularly with Passenger Service Systems (PSS), is preferred.
  • Proficiency in MS Office.
  • Strong knowledge of SQL with experience in database design, optimization, and performance tuning; experience with Oracle Database is an advantage.
  • Solid understanding of distributed systems, system architecture, scalability, high availability, and resilient application design.
  • Proficiency in monitoring, logging, and observability tools such as Prometheus, Grafana, Elasticsearch, or Datadog, used to drive proactive incident detection and reliability strategy.
Work Experience
  • With 6-8 years of experience in Software Engineering, SRE, or Production Support (Java Applications).
  • Hands‑on experience designing, developing, and supporting both microservices and monolithic application architectures.
  • Strong hands‑on experience with Docker and Kubernetes for containerization, orchestration, and production deployment (mandatory).
  • Experience with JBoss Application Server or similar enterprise Java application servers is an added advantage.
  • Proven experience designing and governing CI/CD pipeline standards, GitOps practices, and Git-based workflows across multiple teams.
  • Hands‑on experience with caching technologies (e.g., Redis) and messaging platforms to improve application performance and reliability.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer: Java Microservices & PSS
Lead Site Reliability Engineer: Java Microservices & PSS

ISA • Dubai

On-site
AED 350,000 - 700,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Abu Dhabi

On-site
AED 180,000 - 250,000
System Reliability Engineer (SRE)
System Reliability Engineer (SRE)

TalentRecruit Software Private Limited • Abu Dhabi

On-site
AED 78,000 - 123,000
Supervisor - Maintenance Planning And Technical Records
Supervisor - Maintenance Planning And Technical Records

Air Arabia • Sharjah

On-site
AED 150,000 - 230,000
Technical Services Engineer - Aircraft Maintenance Program
Technical Services Engineer - Aircraft Maintenance Program

Air Arabia • Sharjah

On-site
AED 180,000 - 360,000
Technical Services Engineer - Propulsion
Technical Services Engineer - Propulsion

Air Arabia • Sharjah

On-site
AED 220,000 - 340,000
Site Reliability Engineer - Lead Avrioc Technologies On-site Fast Track available
Site Reliability Engineer - Lead Avrioc Technologies On-site Fast Track available

HireHouse • Abu Dhabi

On-site
AED 300,000 - 600,000
Graduate Engineer Trainee
Graduate Engineer Trainee

Air Arabia • Sharjah

On-site
AED 100,000 - 140,000
Lead Airworthiness Auditor, Air Arabia
Lead Airworthiness Auditor, Air Arabia

Ifairworthy • Sharjah

On-site
AED 223,200 - 334,800
Reliability Engineer - Mechanical
Reliability Engineer - Mechanical

Khazna Data Centers • Dubai

On-site
AED 180,000 - 300,000