Senior Site Reliability Engineer

ISA

Maharashtra

On-site

INR 1,500,000 - 2,300,000

Full time

26 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ISA is seeking an experienced DevOps/SRE-style engineer to ensure stability, reliability, and performance of airline reservation and related systems. You will monitor production environments, perform RCA, and drive automation and CI/CD initiatives using Docker, Kubernetes, and observability tools.

The role requires strong Java/Spring Boot proficiency, knowledge of microservices and monoliths, and hands-on scripting.

Qualifications

  • Bachelor’s Degree in Computer Science, Software Engineering, IT or equivalent.
  • Fluent in English.
  • Strong hands-on Java and Spring Boot expertise.
  • Experience with microservices and monolithic apps.
  • Solid troubleshooting, debugging, and analytical problem-solving.
  • Basic scripting and automation; Python preferred.
  • Proficient in MS Office.
  • Strong SQL with database performance analysis; Oracle a plus.
  • Familiarity with monitoring and observability tools (Prometheus, Grafana, Elasticsearch, Datadog).
  • Exposure to Node.js and Next.js is a bonus.

Responsibilities

  • Troubleshoot and resolve complex production incidents and system performance issues.
  • Perform RCA for failures, outages, and recurring incidents; implement corrective actions.
  • Analyze defects and provide fixes or recommended inputs to development teams.
  • Monitor health, availability, and metrics to ensure reliability and uptime.
  • Improve observability with monitoring, alerting, logging, and dashboards.
  • Support CI/CD pipelines, deployment automation, and release management.
  • Collaborate with development, infrastructure, DevOps, and database teams to improve resiliency.
  • Manage containerized environments using Docker and Kubernetes.
  • Participate in incident management, on-call, and problem management as needed.
  • Review logs, DB performance, and integrations to identify reliability risks.
  • Contribute to automation, runbooks, SOPs, and technical docs.
  • Ensure IT governance, cybersecurity, change management, and standards compliance.
  • Identify opportunities to optimize performance and deployment workflows.

Skills

Java
Spring Boot
Troubleshooting
SQL
Scripting
Automation
Observability
CI/CD
DevOps

Education

Bachelor’s Degree in Computer Science
Information Technology

Tools

Docker
Kubernetes
Jenkins
GitOps
Prometheus
Grafana
Elasticsearch
Datadog
JBoss

Job description

Job Purpose

Responsible for ensuring the stability, reliability, availability, and performance of airline reservation and related operational systems. The role supports production environments through proactive monitoring, troubleshooting, root cause analysis, defect resolution, and continuous improvement of system reliability and operational efficiency. The role also contributes to automation initiatives, CI/CD processes, observability enhancements, and containerized infrastructure management in alignment with business continuity and operational excellence objectives.

Job Purpose

Responsible for ensuring the stability, reliability, availability, and performance of airline reservation and related operational systems. The role supports production environments through proactive monitoring, troubleshooting, root cause analysis, defect resolution, and continuous improvement of system reliability and operational efficiency. The role also contributes to automation initiatives, CI/CD processes, observability enhancements, and containerized infrastructure management in alignment with business continuity and operational excellence objectives.

Key Result Responsibilities
  • Troubleshoot and resolve complex production incidents and system performance issues across airline reservation and associated enterprise applications.
  • Perform detailed root cause analysis (RCA) for application failures, outages, and recurring incidents; ensure corrective and preventive actions are identified and implemented.
  • Analyze defects and either implement fixes directly or provide detailed technical recommendations and inputs to development teams for permanent resolution.
  • Monitor application health, system availability, and operational metrics to ensure service reliability and uptime targets are consistently achieved.
  • Improve observability through enhanced monitoring, alerting, logging, and dashboarding solutions using industry-standard tools and practices.
  • Support and maintain CI/CD pipelines, deployment automation, and release management activities to ensure smooth and reliable software delivery.
  • Work closely with development, infrastructure, DevOps, database, and support teams to improve application resiliency, scalability, and operational efficiency.
  • Support and manage containerized application environments using Docker and Kubernetes.
Key Result Responsibilities-Continued
  • Participate in incident management, on-call support, and problem management activities as required to ensure timely resolution of critical issues.
  • Review application logs, database performance, and system integrations to proactively identify reliability risks and operational bottlenecks.
  • Contribute to automation initiatives, operational runbooks, standard operating procedures, and technical documentation.
  • Ensure compliance with organizational IT governance, cybersecurity, change management, and operational standards.
  • Continuously identify opportunities to optimize application performance, infrastructure utilization, deployment processes, and operational workflows.
Qualifications (Academic, Training, Languages)
  • Bachelor’s Degree in Computer Science, Software Engineering, Information Technology, or equivalent discipline.
  • Fluent in English Language
  • Strong hands‑on expertise in Java and Spring Boot frameworks.
  • Experience working with both microservices architecture and monolithic applications.
  • Strong troubleshooting, debugging, and analytical problem‑solving capabilities.
  • Basic scripting and automation knowledge; Python experience preferred.
  • Proficient in MS Office.
  • Strong SQL knowledge with experience in database troubleshooting and performance analysis; Oracle Database experience is an advantage.
  • Familiarity with monitoring, observability, and logging tools such as Prometheus, Grafana, Elasticsearch, and Datadog.
  • Exposure to Node.js and Next.js is an advantage.
Work Experience
  • 4–7 years of experience supporting Java-based enterprise applications in production environments.
  • Experience with CI/CD tools and deployment pipelines such as Jenkins and GitOps practices.
  • Hands‑on experience with Docker and Kubernetes in enterprise production environments is mandatory.
  • Experience with JBoss application server is considered an advantage. Experience in airline, travel technology, reservation systems, or high‑availability enterprise environments is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

ISA • Maharashtra

On-site
INR 1,200,000 - 1,800,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

ISA • Maharashtra

On-site
INR 1,400,000 - 2,200,000
Technical Architect ( Java & Spring)
Technical Architect ( Java & Spring)

TPF Software Inc. • Tamil Nadu

On-site
INR 3,000,000 - 6,000,000
Application Support Engineer
Application Support Engineer

Common Service Centres (CSC) • Dadri

On-site
INR 450,000 - 650,000
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA Group • Delhi

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]
Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]

SITA • Delhi

On-site
INR 3,500,000 - 7,000,000
Flex Week: work from home up to 2 days
Flex Location: up to 30 days travel
Employee Wellbeing program
+2
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Flex Week: Hybrid
Flex Location: Up to 30 days remote
Employee Wellbeing programs (EAP)
+2
Quality Assurance Engineer - Site Reliability Engineering
Quality Assurance Engineer - Site Reliability Engineering

ISA • Maharashtra

On-site
INR 800,000 - 1,200,000
Application Analyst
Application Analyst

Air Arabia • Pune District

On-site
INR 800,000 - 1,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Saama • Chennai District

On-site
INR 1,200,000 - 1,800,000