Site Reliability Engineer/ Expert/ Specialist

SITA

Delhi

On-site

INR 1,800,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work options
Professional development opportunities
Great Place to Work recognition

Job summary

SITA is seeking an experienced IT Operations professional to ensure high product performance and reliability. You will proactively support services, resolve root causes, and implement improvements that prevent recurrence, focusing on event management, automation, and operational integration.

Responsibilities include managing incidents, defining alerting strategies, and collaborating with Product, Engineering, and Operations to improve service readiness and performance.

Qualifications

  • 5+ years in IT operations, service or infra management.
  • Experience with high-availability systems on Azure/Windows.
  • Proficient in RCA, incident management, and permanent solutions.
  • Hands-on with CI/CD, automation, and IaC.

Responsibilities

  • Build and maintain reliable support systems for high availability and product performance.
  • Manage complex incidents and perform root cause analysis to deliver permanent fixes.
  • Define and maintain event catalogs, alerts, thresholds, and remediation actions.
  • Implement automation for provisioning, monitoring, deployment, self-healing, and recovery.
  • Collaborate with Product, Engineering, Service Architecture, and Operations teams to improve readiness, availability, and performance.
  • Support customer success initiatives through reporting, documentation, and process improvements.
  • Contribute to knowledge management resources and data governance standards.

Skills

RCA and Incident management
CI/CD pipelines
Automation and IaC
Observability focus
Cross-functional collaboration
DevOps practices
Problem management
Data governance

Education

Bachelor's degree in Computer Science / IT / Engineering

Tools

Prometheus
Grafana
Dynatrace
Azure Monitor
AWS CloudWatch
Kubernetes
AKS
Terraform
PowerShell
VMware
Veeam
Veritas

Job description

Overview

At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry. We work with transportation and government clients across locations to deliver fresh solutions and cutting-edge tech that keep operations running smoothly. SITA is recognized as a Great Place to Work by our employees and is committed to empowerment, support, and growth.

Purpose: Ensure high product performance, reliability, and stability by proactively supporting products, resolving root causes of incidents, and implementing improvements that prevent recurrence. This role focuses on event management, automation, service deployment, and operational integration to improve efficiency and collaboration across service operations.

What will you do

Build and maintain reliable support systems to ensure high availability and strong product performance. Manage complex operational cases, incidents, and root cause analysis to deliver permanent fixes. Define and maintain event catalogs, alerts, thresholds, and remediation actions. Implement automation for provisioning, monitoring, deployment, self-healing, and recovery. Collaborate with Product, Engineering, Service Architecture, and Operations teams to improve service readiness, availability, and performance. Support customer success initiatives through reporting, documentation, communication materials, and process improvements. Contribute to knowledge management resources such as FAQs, training materials, and operational guidance. Apply data governance standards, monitor data quality, and act as a subject matter expert for data-related queries.

Qualifications

Experience

  • Bachelor’s degree in computer science, Information Technology, Engineering, or a related field.
  • 5+ years of experience in IT operations, service management, or infrastructure management, including roles such as Site Reliability Engineer, Problem Manager, or DevOps Manager.
  • Proven experience managing high-availability systems and ensuring operational reliability with Azure/Windows Environments.
  • Extensive experience in root cause analysis (RCA), incident management, and developing permanent solutions for recurring service disruptions.
  • Hands-on experience with CI/CD pipelines, automation, system performance monitoring, and infrastructure as code (IaC).
  • Strong collaboration with cross-functional teams to improve operational processes and service delivery.
  • Experience managing deployments, conducting risk assessments, and optimizing event and problem management processes.
  • Familiarity with cloud technologies, containerization, and scalable architectures, including zero-downtime deployment strategies.
  • Development experience.

Operating Systems

  • Strong hands-on expertise with Windows Server (AD, GPO, DNS, DHCP).
  • Solid problem management and troubleshooting skills.
  • Working knowledge of Unix/Linux (Red Hat).
  • PowerShell scripting.

Cloud

  • Strong experience with Azure and AWS.
  • Knowledge and skills in AKS and on-prem Kubernetes.
  • Automation experience, including CI/CD pipelines and exposure to Terraform.

Virtualisation

  • Hands-on experience with VMware (VCF, VCD).

Storage

  • Experience managing SAN and NAS.

Backup

  • Background with Veeam and Veritas backup solutions.

Networking

  • CCNA-level networking knowledge.

Databases

  • Skills in SQL and MongoDB for restore operations and performance tuning.

Observability & Monitoring

  • Experience with enterprise monitoring and observability platforms; hands-on with Prometheus, Grafana, Dynatrace, Azure Monitor, and AWS CloudWatch.
  • Experience in designing proactive monitoring, alerting, and capacity management to improve reliability and reduce incidents.
  • Knowledge of log management, metrics, dashboards, distributed tracing, and root cause analysis.
  • Experience defining and monitoring SLIs, SLOs, and SLAs to support SRE practices.
  • Focus on operational excellence, platform stability, performance optimization, and incident reduction through observability-driven insights.

Core Competencies: (Problem Management)

  • Conduct thorough problem investigations and root cause analyses to diagnose recurring incidents.
  • Coordinate with Incident Management teams and collaborate with PSOs and Engineering/Product teams to implement permanent solutions.
  • Monitor the effectiveness of problem resolution activities and provide regular reporting for continuous improvement.

Qualifications (continued)

  • Bachelor’s degree in computer science, Information Technology, or related field.
  • Minimum 3-5 years of experience in platform and network administration with L2/L3 support on Windows Servers/Azure.
  • Strong problem-solving and analytical skills.
  • Excellent communication and collaboration abilities.
  • Relevant certifications: VMware, AWS/Azure, MCSE, RHCSA/RHCE are highly recommended. CompTIA Security+ or Certified Kubernetes Administrator (AKS or CKA) is desirable. Certifications in cloud platforms (AWS, Azure, Google Cloud) or DevOps methodologies are advantageous.

What We Offer

We value diversity and operate globally. Our offices are inclusive and flexible, with options to work from home where suitable. We provide a supportive environment for development and growth.

Benefits include flexible work options, wellbeing programs, professional development opportunities, and competitive compensation aligned with local markets.

We are an Equal Opportunity Employer and encourage a diverse workforce. We invite applicants to self-identify if they belong to underrepresented groups in the workforce.

Starting Compensation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]
Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]

SITA • Delhi

Hybrid
INR 3,500,000 - 7,000,000
Flex Week: work from home up to 2 days
Flex Location: up to 30 days travel
Employee Wellbeing program
+2
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)

SITA • Delhi

Hybrid
INR 350,000 - 700,000
Associate Infrastructure Engineer
Associate Infrastructure Engineer

SITA • Delhi

Hybrid
INR 1,000,000 - 2,000,000
Flex Week: Work from home up to 2 days/week
Flex Location: Work from any location in the world for up to 30 days a year
Employee Assistance Program for wellbeing
+1
Associate Service Operations Specialist
Associate Service Operations Specialist

SITA Group • Delhi

On-site
INR 600,000 - 900,000
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA Group • Delhi

On-site
INR 1,200,000 - 2,400,000
Associate Service Operations Specialist
Associate Service Operations Specialist

SITA • India

Hybrid
INR 900,000 - 1,200,000
Flex Week: WFH up to 2 days/week
Flex Day
Flex‑Location: up to 30 days/year
+2
Lead Cloud Infrastructure Engineer
Lead Cloud Infrastructure Engineer

SITA • Delhi

Hybrid
INR 2,000,000 - 3,500,000
Hybrid working: up to 2 days WFH per週
Flex-location: up to 30 days remote
Employee wellbeing & EAP
+1
Associate Field Engineer
Associate Field Engineer

SITA Group • NTR

Hybrid
INR 300,000 - 600,000
Flex Week (WFH 2 days)
Flex Day
Flex Location
+3
Associate Field Engineer
Associate Field Engineer

SITA Group • Mumbai

On-site
INR 400,000 - 600,000
Flex Week
Flex Day
Flex-Location
+3
Lead Cloud Infrastructure Engineer
Lead Cloud Infrastructure Engineer

SITA • Bengaluru

Hybrid
INR 2,800,000 - 4,200,000
Hybrid work model
Flexible hours
Flex-location up to 30 days
+3