Lead Site Reliability Engineer/ Expert

SITA Group

Delhi

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SITA is seeking a Reliability Engineer to ensure highly reliable production systems across cloud and on‑prem environments. The role focuses on high availability, disaster recovery readiness, and continuous improvement of service performance, with emphasis on automation, IaC, and robust observability.

You will own recovery strategies, implement self‑healing, manage incidents, and collaborate with cross‑functional teams to deliver provider‑grade reliability for new releases.

Qualifications

  • Proven experience in reliability engineering and scale-out environments.
  • Experience with cloud and on‑prem deployments.

Responsibilities

  • Design & maintain resilient systems ensuring high availability and fault tolerance.
  • Establish DR strategies and improve platform reliability.
  • Drive automation for provisioning, deployment, monitoring and self‑healing.
  • Define SLIs/SLOs and manage incident response and postmortems.
  • Collaborate across product, engineering and service teams to ensure readiness.

Skills

Reliability engineering
DevOps
Observability
Incident management
Cloud platforms

Tools

Kubernetes
Docker
IaC
CI/CD

Job description

Overview
WELCOME TO SITA

At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry.

You’ll find us in 95% of international airports, working closely with over 2,500 transportation and government clients. Each partnership brings unique challenges, and we thrive on delivering fresh solutions and cutting‑edge tech to keep operations running like clockwork. We don’t just move the world forward‑we’re proud to be recognized as a Great Place to Work® by 79% of our employees and certified in most of our growing locations. Here, we feel empowered, supported, and inspired to grow.

Are you ready to love your job?

The adventure begins right here, with you, at SITA.

ABOUT THE ROLE & TEAM

Responsible for ensuring highly reliable, scalable, and resilient production systems across cloud and on‑prem environments. Ensures high availability, disaster recovery readiness, and continuous improvement of service performance. Leads automation initiatives for provisioning, deployment, monitoring, and self‑healing to reduce manual effort and improve stability. Owns the event catalog, operational readiness, and reliability engineering practices to prevent recurrence of incidents and strengthen system resilience. Drives collaboration across Product, Engineering, T&E ICE, and Service Support Architects to ensure provider‑grade reliability and seamless operational integration of new releases.

WHAT YOU’LL DO
  • Reliability Engineering
    • Design & maintain resilient systems ensuring high availability, scalability, and fault tolerance.
    • Ensure effective Disaster Recovery (DR), failover strategies, and resilience engineering across environments.
    • Improve platform reliability, observability, and performance across cloud and on‑premises systems.
    • Establish and maintain SLIs, SLOs, and error budgets to measure and govern service reliability.
    • Take ownership of production availability, capacity planning, performance tuning, and long‑term reliability initiatives.
  • Automation, DevOps & NetOps
    • Drive automation for infrastructure provisioning, deployment, monitoring, and operational workflows.
    • Develop and implement auto‑remediation and self‑healing solutions to reduce manual intervention.
    • Manage CI/CD pipelines and Infrastructure as Code (IaC) frameworks for secure, repeatable deployments.
    • Implement and manage zero‑downtime deployment strategies (blue‑green, canary, rolling).
    • Support containerized and cloud‑native platforms including Kubernetes, Docker, and distributed systems.
    • Support NetOps tooling and network observability, ensuring visibility into network performance, events, and operational health.
  • Incident, Problem & Event Management
    • Perform incident management, production troubleshooting, and lead RCA/PMIR (Postmortem) for critical outages.
    • Proactively identify reliability gaps, performance bottlenecks, and operational risks.
    • Optimize incident, event, and problem management processes to reduce MTTR and improve operational efficiency.
    • Define and maintain the event catalog, thresholds, and remediation workflows.
    • Develop event response protocols and ensure teams are trained for rapid incident handling.
  • Observability & Monitoring
    • Build and maintain observability solutions using monitoring, logging, tracing, and alerting platforms.
    • Implement APM, distributed tracing, and proactive alerting to detect issues early.
    • Integrate network telemetry and NetOps monitoring tools into the overall observability stack.
    • Collaborate with stakeholders to improve event coverage and post‑event learning.
    • Experience with AI‑assisted observability, anomaly detection, and predictive alerting.
  • Deployment & Operational Readiness
    • Own the quality of new release deployments for the PSO.
    • Conduct operational readiness assessments and manage deployment risk.
    • Ensure supportability for new applications, platform releases, and infrastructure changes.
    • Coordinate with internal/external stakeholders to drive continuous service improvement.
  • Cross‑Functional Collaboration
    • Work closely with Development, Platform Engineering, Product, T&E ICE, and Service Support Architects to embed reliability best practices.
    • Collaborate with vendors and engineering teams to enhance system reliability and operational excellence.
    • Support new product productization as SGS technical expert and ensure operational readiness.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)

SITA • Delhi

Hybrid
INR 350,000 - 700,000
Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]
Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]

SITA • Delhi

Hybrid
INR 3,500,000 - 7,000,000
Flex Week: work from home up to 2 days
Flex Location: up to 30 days travel
Employee Wellbeing program
+2
Site Reliability Engineer/ Expert/ Specialist
Site Reliability Engineer/ Expert/ Specialist

SITA • Delhi

On-site
INR 1,800,000 - 2,400,000
Flexible work options
Professional development opportunities
Great Place to Work recognition
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

TTEC Digital • Hyderabad

On-site
INR 4,500,000 - 6,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Technologies Pvt. Ltd. • Pune District

On-site
INR 900,000 - 1,400,000
Engineering Manager
Engineering Manager

WaferWire Cloud Technologies • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Bengaluru

On-site
INR 2,400,000 - 3,400,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Site Reliability Specialist - High Availability
Site Reliability Specialist - High Availability

Freelanceshop • Gwalior District

Hybrid
INR 1,200,000 - 2,000,000
Competitive salary
Health and life insurance
Flexible work arrangements
+2