Lead Site Reliability Engineer/ Expert

SITA

Bengaluru

Hybrid

INR 3,500,000 - 6,000,000

Full time

36 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Flex Week: Hybrid
Flex Location: Up to 30 days remote
Employee Wellbeing programs (EAP)
Professional Development: LinkedIn You
Competitive Benefits

Job summary

SITA is hiring a hands-on Lead Site Reliability Engineer (SRE) to ensure reliability and observability of production systems. You will lead incident investigations, drive permanent fixes, and collaborate with Development, Product, and Operations teams.

Strong troubleshooting across Kubernetes and production environments is essential. The role requires deep expertise in RCA, scripting, and CI/CD, with a focus on reducing outages and improving resilience.

Qualifications

  • 8+ years’ experience in SRE, DevOps, or production engineering for high-availability systems.
  • Strong RCA and permanent issue resolution skills.
  • Hands-on production troubleshooting using logs, metrics, and traces.
  • Experience with Java and/or .NET applications in production.
  • Hands-on experience with Kubernetes and containerized workloads.

Responsibilities

  • Analyze production incidents using logs, metrics, and traces to identify impacted application code and execution paths.
  • Diagnose system issues and determine whether root cause is related to application code, configuration, Kubernetes, or infrastructure.
  • Troubleshoot Kubernetes workloads, including runtime behavior, networking, health probes, and failure scenarios.
  • Serve as the technical escalation point during critical incidents, providing clear and timely guidance.
  • Lead root cause analysis (RCA) and drive permanent corrective actions to improve reliability.
  • Enhance observability and alerting to improve issue detection and resolution.
  • Partner with Development and Platform teams to resolve systemic issues and strengthen service reliability.
  • Automate repetitive operational tasks and promote engineering best practices.
  • Assess the impact of deployments on production environments through CI/CD pipeline expertise.
  • Monitor system performance and reliability, identifying opportunities to improve resilience and reduce outages.

Skills

Root cause analysis
Troubleshooting
Java/.NET production apps
Scripting (Python/Bash)
Communication

Education

Bachelor’s degree in Computer Science or related field

Tools

Kubernetes
Azure DevOps
Jenkins
GitHub Actions
Monitoring & observability tools

Job description

! At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry.

Overview

Welcome to SITA! At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry. You’ll find us in 95% of international airports, working closely with over 2,500 transportation and government clients. Each partnership brings unique challenges, and we thrive on delivering fresh solutions and cutting‑edge tech to keep operations running like clockwork. We don’t just move the world forward—we’re proud to be recognized as a Great Place to Work® by 79% of our employees and certified in most of our growing locations. Here, we feel empowered, supported, and inspired to grow. Are you ready to love your job? The adventure begins right here, with you, at SITA.

About The Role & Team

We are seeking a hands‑on Lead Site Reliability Engineer (SRE) with strong expertise across application support, Kubernetes environments, and CI/CD pipelines. This role is responsible for ensuring the reliability, performance, and observability of production systems through deep technical analysis and proactive engineering. The successful candidate will lead incident and problem investigations, identify root causes, and drive permanent resolutions in close collaboration with Development, Product, and Operations teams. This is an engineering‑focused role requiring strong troubleshooting skills, production support experience, and a commitment to continuous service improvement.

What You Will Do
  • Analyze production incidents using logs, metrics, and traces to identify impacted application code and execution paths.
  • Diagnose system issues and determine whether the root cause is related to application code, configuration, Kubernetes, or infrastructure.
  • Troubleshoot Kubernetes workloads, including runtime behavior, networking, health probes, and failure scenarios.
  • Serve as the technical escalation point during critical incidents, providing clear and timely guidance.
  • Lead root cause analysis (RCA) and drive permanent corrective actions to improve reliability.
  • Enhance observability and alerting to improve issue detection and resolution.
  • Partner with Development and Platform teams to resolve systemic issues and strengthen service reliability.
  • Automate repetitive operational tasks and promote engineering best practices.
  • Assess the impact of deployments on production environments through CI/CD pipeline expertise.
  • Monitor system performance and reliability, identifying opportunities to improve resilience and reduce outages.
Qualifications
About Your Skills
  • Min 8 years’ experience in SRE, DevOps, or Production Engineering supporting high‑availability systems.
  • Strong expertise in root cause analysis (RCA) and permanent issue resolution.
  • Hands‑on troubleshooting of production environments using logs, metrics, and traces.
  • Strong knowledge of Java and/or .NET applications in production.
  • Hands‑on experience with Kubernetes and containerized workloads.
  • Experience with monitoring and observability tools for distributed systems.
  • Familiarity with CI/CD pipelines and deployment tools (e.g., Azure DevOps, Jenkins, GitHub Actions).
  • Scripting and automation skills using Python, Bash, or similar.
  • Strong analytical, communication, and cross‑functional collaboration skills.
  • Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent experience.
What We Offer

We value diversity, operating in 200 countries and spanning 60 languages and cultures. Our inclusive offices are comfortable and fun, with the flexibility to work from home. Join our team and step closer to your best life.

  • Flex Week: Hybrid (2 days from home and 3 days in the office.)
  • Flex Location: Take up to 30 days a year to work from any location in the world.
  • Employee Wellbeing: We’ve got you covered with our Employee Assistance Program (EAP), for you and your dependents 24/7, 365 days/year. We also offer Champion Health a personalized platform that supports a range of well‑being needs.
  • Professional Development: Level up your skills with our training platforms, including LinkedIn Learning!
  • Competitive Benefits: Competitive benefits that make sense with both your local market and employment status.

SITA is an Employment Equity Employer and values a diverse workforce. In support of our Employment Equity Program, women, Aboriginal people, members of visible minorities, and/or persons with disabilities are encouraged to apply and self‑identify in the application process.

Salary / Compensation Note

Hidden (-999)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]
Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]

SITA • Delhi

On-site
INR 3,500,000 - 7,000,000
Flex Week: work from home up to 2 days
Flex Location: up to 30 days travel
Employee Wellbeing program
+2
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA • Delhi

Hybrid
INR 4,000,000 - 7,000,000
Flex Week
Flex Day
Flex-Location
+3
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)

SITA • Delhi

Hybrid
INR 350,000 - 700,000
Site Reliability Engineer/ Expert/ Specialist
Site Reliability Engineer/ Expert/ Specialist

SITA • Delhi

On-site
INR 1,800,000 - 2,400,000
Flexible work options
Professional development opportunities
Great Place to Work recognition
Site Reliability Engineer - Vice President
Site Reliability Engineer - Vice President

Citi • Maharashtra

On-site
INR 4,000,000 - 7,000,000
Senior Infrastructure Engineer(Azure)
Senior Infrastructure Engineer(Azure)

SITA • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Flex Week
Flex Day
Flex Location
+2
Lead Software Developer(Java)
Lead Software Developer(Java)

SITA • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Flex Week: work from home up to 2 days
Flex Location: global work options
Employee Wellbeing
+2
Senior Infrastructure Engineer
Senior Infrastructure Engineer

SITA • Delhi

Hybrid
INR 1,200,000 - 1,800,000
Flex Week: Work from home up to 2 days
Flex Day
Flex-Location: up to 30 days/year
+3
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA Group • Delhi

On-site
INR 1,200,000 - 2,400,000
Senior Infrastructure Engineer(Azure)
Senior Infrastructure Engineer(Azure)

SITA Group • Bengaluru

Hybrid
INR 1,800,000 - 3,200,000
Flex Week: Work from home up to 2 days
Flex Day: Flexible work hours
Flex Location: Up to 30 days remote
+2