Lead Site Reliability Engineer/ Expert

Sita

New Delhi

On-site

INR 2,600,000 - 4,200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Sita is seeking an experienced Reliability Engineer to ensure highly available, scalable production systems across cloud and on-prem environments. You will own the event catalog, operational readiness, and R&E practices to prevent incidents and strengthen resilience.

You will lead automation for provisioning, deployment, monitoring, and self-healing, coordinating with Product, Engineering, and Service Support Architects.

Qualifications

  • 8+ years in IT operations, service management, or infrastructure reliability.
  • Strong experience with high availability, resilience engineering and DR readiness.
  • Hands-on CI/CD, automation, IaC, and self-healing workflows.
  • Experience with observability platforms and NetOps monitoring.

Responsibilities

  • Design and maintain resilient systems with high availability, scalability and fault tolerance.
  • Ensure DR readiness and failover strategies across environments.
  • Improve platform reliability, observability and performance for cloud/on-prem.
  • Define and govern SLIs/SLOs and error budgets for reliability.
  • Own production availability, capacity planning and performance tuning.
  • Drive automation for provisioning, deployment, monitoring and workflows.
  • Develop auto-remediation and self-healing solutions.
  • Manage CI/CD pipelines and IaC frameworks for secure deployments.
  • Implement zero-downtime deployment strategies (blue-green, canary).
  • Support Kubernetes, Docker and distributed systems in production.
  • Support NetOps tooling and network observability for health.
  • Lead incident, problem, and postmortem activities to improve resilience.
  • Collaborate across teams to embed reliability best practices.

Skills

SRE practices
Incident management
CI/CD
Automation
Observability
Kubernetes
Zero-downtime
NetOps monitoring

Education

Bachelor's degree in CS/IT/Engineering
Certifications: ITIL, CCNP/CCIE, Palo Alto, SASE, SD-WAN
Cloud/DevOps certifications (AWS, Azure, GCP)
Automation/IaC certifications (Ansible, Terraform)
Observability tools certifications (Dynatrace, Prometheus, Grafana, ELK)
ServiceNow/Jira operational tooling certifications

Tools

Terraform
Ansible
Dynatrace
Prometheus
Grafana
ELK
Kubernetes

Job description

Job Summary

Responsible for ensuring highly reliable, scalable, and resilient production systems across cloud and on-prem environments. Ensures high availability, disaster recovery readiness, and continuous improvement of service performance. Leads automation initiatives for provisioning, deployment, monitoring, and self-healing to reduce manual effort and improve stability. Owns the event catalog, operational readiness, and reliability engineering practices to prevent recurrence of incidents and strengthen system resilience. Drives collaboration across Product, Engineering, T&E ICE, and Service Support Architects to ensure provider-grade reliability and seamless operational integration of new releases.

Responsibilities
  • Reliability Engineering: Design & maintain resilient systems ensuring high availability, scalability, and fault tolerance.
  • Reliability Engineering: Ensure effective Disaster Recovery (DR), failover strategies, and resilience engineering across environments.
  • Reliability Engineering: Improve platform reliability, observability, and performance across cloud and on-premises systems.
  • Reliability Engineering: Establish and maintain SLIs, SLOs, and error budgets to measure and govern service reliability.
  • Reliability Engineering: Take ownership of production availability, capacity planning, performance tuning, and long-term reliability initiatives.
  • Automation, DevOps & NetOps: Drive automation for infrastructure provisioning, deployment, monitoring, and operational workflows.
  • Automation, DevOps & NetOps: Develop and implement auto-remediation and self-healing solutions to reduce manual intervention.
  • Automation, DevOps & NetOps: Manage CI/CD pipelines and Infrastructure as Code (IaC) frameworks for secure, repeatable deployments.
  • Automation, DevOps & NetOps: Implement and manage zero-downtime deployment strategies (blue-green, canary, rolling).
  • Automation, DevOps & NetOps: Support containerized and cloud-native platforms including Kubernetes, Docker, and distributed systems.
  • Automation, DevOps & NetOps: Support NetOps tooling and network observability, ensuring visibility into network performance, events, and operational health.
  • Incident, Problem & Event Management: Perform incident management, production troubleshooting, and lead RCA/PMIR (Postmortem) for critical outages.
  • Incident, Problem & Event Management: Proactively identify reliability gaps, performance bottlenecks, and operational risks.
  • Incident, Problem & Event Management: Optimize incident, event, and problem management processes to reduce MTTR and improve operational efficiency.
  • Incident, Problem & Event Management: Define and maintain the event catalog, thresholds, and remediation workflows.
  • Incident, Problem & Event Management: Develop event response protocols and ensure teams are trained for rapid incident handling.
  • Observability & Monitoring: Build and maintain observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Observability & Monitoring: Implement APM, distributed tracing, and proactive alerting to detect issues early.
  • Observability & Monitoring: Integrate network telemetry and NetOps monitoring tools into the overall observability stack.
  • Observability & Monitoring: Collaborate with stakeholders to improve event coverage and post-event learning.
  • Observability & Monitoring: Experience with AI-assisted observability, anomaly detection, and predictive alerting.
  • Deployment & Operational Readiness: Own the quality of new release deployments for the PSO.
  • Deployment & Operational Readiness: Conduct operational readiness assessments and manage deployment risk.
  • Deployment & Operational Readiness: Ensure supportability for new applications, platform releases, and infrastructure changes.
  • Deployment & Operational Readiness: Coordinate with internal/external stakeholders to drive continuous service improvement.
  • Cross-Functional Collaboration: Work closely with Development, Platform Engineering, Product, T&E ICE, and Service Support Architects to embed reliability best practices.
  • Cross-Functional Collaboration: Collaborate with vendors and engineering teams to enhance system reliability and operational excellence.
  • Cross-Functional Collaboration: Support new product productization as SGS technical expert and ensure operational readiness.
Who you are
Education and Professional Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field. Master's degree preferred for senior roles.
  • Relevant certifications such as ITIL, CCNP/CCIE, Palo Alto Security, SASE, SDWAN, Juniper Mist/Aruba, CompTIA Security+, or Certified Kubernetes Administrator (CKA).
  • Certifications in cloud platforms (AWS, Azure, Google Cloud) or DevOps methodologies.
  • Certifications in automation and IaC tools (Ansible, Terraform).
  • Certifications in observability and monitoring platforms (Dynatrace, Prometheus, Grafana, ELK).
  • Certifications in ServiceNow, Jira, or other operational tooling.
Experience
  • 8+ years in IT operations, service management, or infrastructure reliability, including roles such as Site Reliability Engineer, Problem Manager, or DevOps Engineer.
  • Strong experience with high availability systems, resilience engineering, and DR readiness.
  • Deep expertise in RCA, incident management, PMIR, and implementing permanent fixes for recurring issues.
  • Hands on experience with CI/CD, automation, IaC, and self healing/auto remediation workflows.
  • Proficiency in observability platforms (APM, logging, tracing, alerting) and integrating network telemetry / NetOps monitoring.
  • Experience defining and governing SLIs, SLOs, and error budgets to improve service reliability.
  • Experience with Kubernetes, containerized workloads, and distributed systems.
  • Experience managing deployments, operational readiness, risk assessments, and improving event/problem management processes.
  • Strong cross functional collaboration with Development, Operations, Engineering, Product, T&E ICE, and SSA.
  • Familiarity with cloud platforms, scalable architectures, and zero downtime deployment strategies.
Technical Skills
  • Cloud Infrastructure - AWS/Azure, Linux, virtualization, HA/DR architecture.
  • Automation & IaC - Ansible, Terraform, CI/CD pipelines, self-healing workflows.
  • Observability & Monitoring - APM, logging, tracing, alerting, Dynatrace, Prometheus, Grafana, ELK.
  • NetOps Monitoring - network telemetry, event monitoring, and operational visibility tools.
  • Containerization & Orchestration - Docker, Kubernetes, distributed systems.
  • Deployment & Release Engineering - zero-downtime strategies (blue-green, canary), operational readiness.
  • Programming & Scripting - Python, Bash, PowerShell for automation and tooling.
  • Reliability Engineering - SLIs/SLOs, error budgets, capacity planning, performance tuning.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA Group • Delhi

On-site
INR 1,200,000 - 2,400,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Hackajob • Ahmedabad District, Gurugram District, Mumbai

On-site
INR 1,800,000 - 2,600,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Zeta Global • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)

SITA • Delhi

Hybrid
INR 350,000 - 700,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000