Senior Site Reliability Engineer

ISA

Sharjah

On-site

AED 240,000 - 360,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ISA is seeking an experienced DevOps/Platform Engineer to ensure stability, reliability, and performance of airline reservation systems. You will contribute to proactive monitoring, RCA, and automation initiatives, while supporting CI/CD, observability, and containerized infrastructure in a production environment.

The role requires hands-on Java/Spring Boot experience, expertise in microservices and cloud-native patterns, and strong scripting, SQL, and diagnostics capabilities.

Qualifications

  • Bachelor's degree in CS/Engineering/IT or equivalent.
  • Fluent English required.
  • Strong hands-on Java and Spring Boot experience.
  • Experience with microservices and monolithic apps.
  • Strong troubleshooting and analytical skills.
  • Scripting and automation basics; Python preferred.
  • SQL knowledge with performance analysis; Oracle a bonus.
  • Familiarity with observability tools like Prometheus, Grafana, Elasticsearch, Datadog.

Responsibilities

  • Troubleshoot and resolve complex production incidents and system performance issues across airline reservation and related applications.
  • Perform root cause analysis for failures and incidents; implement corrective actions.
  • Analyze defects and provide fixes or guidance to development teams for permanent resolution.
  • Monitor health and availability to meet reliability and uptime targets.
  • Improve observability with monitoring, alerting, logging, and dashboards.
  • Support CI/CD pipelines, deployment automation, and release management.
  • Collaborate with Dev, Infra, DevOps, DB, and Support teams to improve resiliency and efficiency.
  • Manage containerized environments using Docker and Kubernetes.
  • Participate in incident management, on-call support, and problem management as needed.
  • Review logs, DB performance, and integrations to identify reliability risks.
  • Contribute to automation, runbooks, SOPs, and technical docs.
  • Ensure compliance with IT governance, cybersecurity, and change management.
  • Identify optimization opportunities for performance and deployment workflows.

Skills

Java
Spring Boot
Microservices
Fluent English
Troubleshooting
Python
SQL
Oracle
Docker
Kubernetes
Jenkins
GitOps

Education

Bachelor's Degree in Computer Science, Software Engineering, Information Technology, or equivalent discipline

Tools

Prometheus
Grafana
Elasticsearch
Datadog

Job description

Job Purpose

Responsible for ensuring the stability, reliability, availability, and performance of airline reservation and related operational systems. The role supports production environments through proactive monitoring, troubleshooting, root cause analysis, defect resolution, and continuous improvement of system reliability and operational efficiency. The role also contributes to automation initiatives, CI/CD processes, observability enhancements, and containerized infrastructure management in alignment with business continuity and operational excellence objectives.

Key Result Responsibilities
  • Troubleshoot and resolve complex production incidents and system performance issues across airline reservation and associated enterprise applications.
  • Perform detailed root cause analysis (RCA) for application failures, outages, and recurring incidents; ensure corrective and preventive actions are identified and implemented.
  • Analyze defects and either implement fixes directly or provide detailed technical recommendations and inputs to development teams for permanent resolution.
  • Monitor application health, system availability, and operational metrics to ensure service reliability and uptime targets are consistently achieved.
  • Improve observability through enhanced monitoring, alerting, logging, and dashboarding solutions using industry-standard tools and practices.
  • Support and maintain CI/CD pipelines, deployment automation, and release management activities to ensure smooth and reliable software delivery.
  • Work closely with development, infrastructure, DevOps, database, and support teams to improve application resiliency, scalability, and operational efficiency.
  • Support and manage containerized application environments using Docker and Kubernetes.
Key Result Responsibilities-Continued
  • Participate in incident management, on-call support, and problem management activities as required to ensure timely resolution of critical issues.
  • Review application logs, database performance, and system integrations to proactively identify reliability risks and operational bottlenecks.
  • Contribute to automation initiatives, operational runbooks, standard operating procedures, and technical documentation.
  • Ensure compliance with organizational IT governance, cybersecurity, change management, and operational standards.
  • Continuously identify opportunities to optimize application performance, infrastructure utilization, deployment processes, and operational workflows.
Qualifications (Academic, Training, Languages)
  • Bachelor's Degree in Computer Science, Software Engineering, Information Technology, or equivalent discipline.
  • Fluent in English Language
  • Strong hands‑on expertise in Java and Spring Boot frameworks.
  • Experience working with both microservices architecture and monolithic applications.
  • Strong troubleshooting, debugging, and analytical problem‑solving capabilities.
  • Basic scripting and automation knowledge; Python experience preferred.
  • Proficient in MS Office.
  • Strong SQL knowledge with experience in database troubleshooting and performance analysis; Oracle Database experience is an advantage.
  • Familiarity with monitoring, observability, and logging tools such as Prometheus, Grafana, Elasticsearch, and Datadog.
Work Experience
  • 4-7 years of experience supporting Java‑based enterprise applications in production environments.
  • Experience with CI/CD tools and deployment pipelines such as Jenkins and GitOps practices.
  • Hands‑on experience with Docker and Kubernetes in enterprise production environments is mandatory.
  • Experience with JBoss application server is considered an advantage. Experience in airline, travel technology, reservation systems, or high‑availability enterprise environments is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Air Arabia • Sharjah

On-site
AED 223,000 - 402,000
Senior SRE – Airline Reservations, Observability & CI/CD
Senior SRE – Airline Reservations, Observability & CI/CD

ISA • Sharjah

On-site
AED 240,000 - 360,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Dubai

On-site
AED 200,000 - 300,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Abu Dhabi

On-site
AED 180,000 - 250,000
Application Support Specialist
Application Support Specialist

Abu Dhabi Aviation • Abu Dhabi Emirate

On-site
AED 180,000 - 240,000
Analyst - IT Solutions (Travel Domain)
Analyst - IT Solutions (Travel Domain)

ISA • Sharjah

On-site
AED 279,000 - 469,000
Senior SRE - Airline Tech & Observability
Senior SRE - Airline Tech & Observability

Air Arabia • Sharjah

On-site
AED 223,000 - 402,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Open Innovation AI • Abu Dhabi Emirate

On-site
AED 150,000 - 210,000
Application Support Engineer (L2 – Enterprise Applications)
Application Support Engineer (L2 – Enterprise Applications)

eMinds • Abu Dhabi Emirate

On-site
AED 201,000 - 312,000
Analyst - IT Solutions Travel Domain
Analyst - IT Solutions Travel Domain

TALENTMATE • Sharjah

On-site
AED 201,000 - 312,000