Lead Site Reliability Engineer

ISA

Maharashtra

On-site

INR 1,400,000 - 2,200,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

ISA in Maharashtra, India seeks an experienced Site Reliability Engineer to lead reliability for mission-critical Java-based applications and airline reservation systems, applying SRE best practices.

You will own SLAs/SLOs, incident management, and capacity planning while championing CI/CD, GitOps, and containerization with Docker and Kubernetes. 6–8 years in SRE/Production Support (Java) are required, with airline domain experience preferred.

Qualifications

  • Bachelor’s degree in Computer Engineering/CS/IT.
  • Fluent in English.
  • Strong Java, Spring Boot expertise; enterprise issue troubleshooting.
  • Experience with airline/aviation systems is preferred.
  • Knowledge of SQL and Oracle DB is advantageous.
  • Experience with monitoring/observability tools.

Responsibilities

  • Lead SRE initiatives for Java microservices and monolithic apps.
  • Define and manage SLAs, SLOs, and error budgets.
  • Lead RCAs for production incidents and implement preventive actions.
  • Drive architecture improvements for performance, scalability, and availability.
  • Mentor teammates on incident management and operational excellence.
  • Collaborate with engineering, infra, and cross-functional teams for production readiness.
  • Define CI/CD and GitOps standards; own containerization strategy (Docker/Kubernetes).
  • Own on-call framework, post-incident reviews, and reliability reporting.
  • Evaluate new tools for reliability and capacity planning.

Skills

Java
Spring Boot
English fluency
Incidents & RCA
CI/CD
GitOps

Education

Bachelor's Degree in Computer Science/Engineering/IT

Tools

Docker
Kubernetes
JBoss
Prometheus
Grafana
Elasticsearch
Datadog
Redis
Oracle Database
Node.js
Next.js

Job description

Job Purpose

To lead the reliability, availability, performance, and continuous improvement of Air Arabia's mission-critical Java-based applications and airline reservation systems by implementing Site Reliability Engineering (SRE) best practices. Responsible for driving operational excellence through proactive monitoring, automation, incident management, scalability improvements, and collaboration with cross-functional teams, while ensuring compliance with organizational policies, industry standards, and applicable regulatory requirements.



Key Result Responsibilities


  • Lead Site Reliability Engineering (SRE) initiatives for Java-based microservices and monolithic applications supporting mission-critical airline operations and Passenger Service Systems (PSS).

  • Establish, monitor, and continuously improve system reliability by defining and managing Service Level Agreements (SLAs), Service Level Objectives (SLOs), and error budgets.

  • Lead and facilitate root cause analysis (RCA) for complex production incidents, ensuring timely resolution and implementation of preventive measures to minimize recurrence.

  • Drive architectural enhancements to improve system performance, scalability, resilience, availability, and operational efficiency through optimization techniques such as caching and distributed system design.

  • Mentor and provide technical guidance to team members by promoting best practices in incident management, production support, troubleshooting, and operational excellence.



Key Result Responsibilities-Continued


  • Collaborate closely with software engineering, infrastructure, and cross-functional teams to enhance application design, improve system reliability, and ensure production readiness.

  • Own and define CI/CD pipeline standards, GitOps practices, and infrastructure automation strategy, setting the framework that engineers across the team execute against.

  • Set the containerization and orchestration strategy (Docker, Kubernetes) for the team, defining standards for scalability, resilience, and high availability that other engineers implement.

  • Own the on-call escalation framework, ensuring adequate coverage and clear escalation paths, and lead post-incident reviews.

  • Evaluate new tools and technologies to strengthen reliability and represent SRE in capacity planning and release readiness reviews.

  • Own reliability reporting to management, translating SLA/SLO performance and incident trends into clear business updates.



Qualifications (Academic, Training, Languages)


  • Bachelor’s Degree in Computer Engineering/ Computer Science/ Information Technology.

  • Fluent in English Language.

  • Strong expertise in Java, Spring Boot, and troubleshooting complex production issues within enterprise environments.

  • Airline, aviation, or travel industry experience, particularly with Passenger Service Systems (PSS), is preferred.

  • Proficiency in MS Office.

  • Strong knowledge of SQL with experience in database design, optimization, and performance tuning; experience with Oracle Database is an advantage.

  • Solid understanding of distributed systems, system architecture, scalability, high availability, and resilient application design.

  • Proficiency in monitoring, logging, and observability tools such as Prometheus, Grafana, Elasticsearch, or Datadog, used to drive proactive incident detection and reliability strategy.

  • Exposure to Node.js and Next.js is an advantage



Work Experience


  • With 6-8 years of experience in Software Engineering, SRE, or Production Support (Java Applications).

  • Hands-on experience designing, developing, and supporting both microservices and monolithic application architectures.

  • Strong hands-on experience with Docker and Kubernetes for containerization, orchestration, and production deployment (mandatory).

  • Experience with JBoss Application Server or similar enterprise Java application servers is an added advantage.

  • Proven experience designing and governing CI/CD pipeline standards, GitOps practices, and Git-based workflows across multiple teams.

  • Hands-on experience with caching technologies (e.g., Redis) and messaging platforms to improve application performance and reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

ISA • Maharashtra

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ISA • Maharashtra

On-site
INR 1,500,000 - 2,300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
SRE Lead
SRE Lead

3across • Bengaluru

Hybrid
INR 1,500,000 - 2,300,000
Site Reliability Engineer
Site Reliability Engineer

Peoplefy • Pune District

On-site
INR 1,200,000 - 1,800,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineering Lead (Application SRE Lead)
Site Reliability Engineering Lead (Application SRE Lead)

Hirexa Solutions • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

Infosys • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Flex Week: Hybrid
Flex Location: Up to 30 days remote
Employee Wellbeing programs (EAP)
+2
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Disney Experiences • Bengaluru

On-site
INR 5,500,000 - 7,500,000