Site Reliability Engineering (SRE) Lead

Sidglobal

Hyderabad

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SID Global Solutions is seeking an experienced Site Reliability Engineering Lead to own cloud infrastructure, Kubernetes, API management, and production operations for enterprise applications. You will drive reliability, observability, incident response, and automation to ensure service availability.

You will collaborate with Development, DevOps, Infrastructure, Platform Engineering, and Support teams to maintain highly available environments, reduce risks, and improve platform reliability

Qualifications

  • 7–10 years of SRE/DevOps/Platform Engineering or Production Support experience.
  • Hands-on experience with Kubernetes, GCP, and Google Apigee.
  • Strong understanding of distributed systems and cloud-native applications.
  • Experience in Incident, Problem, Change, and Release Management.
  • Ability to coordinate cross-functional teams during critical incidents.

Responsibilities

  • Define and implement SLIs, SLOs, and error budgets to improve platform reliability.
  • Review architectures to ensure scalability, resilience, and high availability.
  • Drive reliability improvements for cloud-native apps and distributed systems.
  • Lead Major Incident Management (P1/P2) and act as escalation point for production issues.
  • Coordinate with cross-functional teams to restore services within SLAs.
  • Develop automation scripts and self-healing capabilities to reduce toil.

Skills

Google Cloud Platform (GCP)
Google Apigee API Management
NGINX/Load balancing
Linux/Unix Administration
Python/Bash/Go scripting
CI/CD pipelines
Terraform/IaC

Tools

Kubernetes

Job description

Site Reliability Engineering (SRE) Lead

Employment Type: Full-Time

About SID Global Solutions

SID Global Solutions (SIDGS) is a leading Digital Engineering and Technology Services company specializing in Cloud, API Management, DevOps, Platform Engineering, Kubernetes, Microservices, and Digital Transformation. We are looking for an experienced Site Reliability Engineering (SRE) Lead to drive platform reliability, operational excellence, and production stability for mission‑critical enterprise applications.

Job Summary

We are seeking a highly skilled SRE Lead with strong expertise in cloud infrastructure, Kubernetes, API Management, and production operations. The ideal candidate will be responsible for ensuring the availability, scalability, performance, and reliability of enterprise applications while leading incident management, observability, and automation initiatives.

In this role, you will work closely with Development, DevOps, Infrastructure, Platform Engineering, and Application Support teams to maintain highly available production environments, reduce operational risks, and improve service reliability through automation and proactive monitoring.

Key Responsibilities
Reliability Engineering
  • Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to improve platform reliability.
  • Review application architecture and infrastructure designs to ensure scalability, resilience, and high availability.
  • Drive reliability improvements across cloud‑native applications and distributed systems.
  • Identify opportunities to eliminate operational bottlenecks through automation and process improvements.
  • Lead Major Incident Management (P1/P2) activities and act as the primary technical escalation point for critical production issues.
  • Coordinate with cross‑functional teams to restore services within defined SLAs.
  • Conduct Root Cause Analysis (RCA) and drive preventive and corrective actions to reduce recurring incidents.
  • Participate in Change Management and Release activities to ensure production stability.
  • Prepare incident reports and communicate status updates to business and technology stakeholders.
  • Develop automation scripts and self‑healing solutions to improve operational efficiency.
  • Build reusable operational runbooks and standard operating procedures.
  • Automate routine operational tasks using Python, Bash, or similar scripting languages.
  • Support Infrastructure as Code (IaC) initiatives and CI/CD pipeline improvements.
Monitoring & Observability
  • Design and maintain enterprise monitoring and alerting solutions.
  • Created dashboards and alerts to proactively monitor application health, infrastructure, APIs, and Kubernetes environments.
  • Analyze performance trends and recommend improvements for system stability and capacity planning.
  • Manage and troubleshoot Google Apigee API Gateway configurations, policies, and traffic routing.
  • Monitor Kubernetes clusters, workloads, namespaces, ingress controllers, and container health.
  • Optimize application performance and resource utilization within Kubernetes environments.
  • Support production deployments and post‑release validation activities.
  • Mentor SRE, DevOps, and Production Support engineers.
  • Establish operational best practices, troubleshooting guidelines, and technical documentation.
  • Collaborate with Development, QA, Infrastructure, Security, and Business teams to improve platform reliability.
  • Drive a culture of continuous improvement, automation, and operational excellence.
Required Technical Skills
  • Strong experience with Google Cloud Platform (GCP).
  • Hands‑on expertise in Google Apigee API Management.
  • Good understanding of NGINX, reverse proxy, and load balancing concepts.
  • Strong knowledge of Linux/Unix Administration.
  • Experience with Python, Bash, or Go scripting.
  • Familiarity with CI/CD pipelines, Git, and Jenkins.
  • Understanding of Infrastructure as Code (Terraform or equivalent is preferred).
Monitoring & Observability Tools

Experience with one or more of the following:

  • Dynatrace
  • Grafana
  • Splunk
  • ELK Stack
  • AppDynamics
Required Qualifications
  • 7–10 years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, or Production Support.
  • Strong experience supporting enterprise production environments.
  • Hands‑on experience with Kubernetes, GCP, and Google Apigee.
  • Good understanding of distributed systems, microservices architecture, and cloud‑native applications.
  • Experience in Incident, Problem, Change, and Release Management.
  • Ability to troubleshoot complex production issues and coordinate cross‑functional teams during critical incidents.
Preferred Qualifications
  • Experience in Banking, Financial Services, or other enterprise environments.
  • Google Cloud Professional Certification.
  • Kubernetes Certification (CKA/CKAD) is an added advantage.
  • Experience with Service Mesh technologies such as Istio is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE) Lead
Site Reliability Engineering (SRE) Lead

SID Global Solutions • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer (SRE) – GCP Platform
Site Reliability Engineer (SRE) – GCP Platform

ITC Infotech • Bengaluru

On-site
INR 900,000 - 1,300,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Synechron • Bengaluru, Hyderabad

Hybrid
INR 4,200,000 - 6,300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Atria Convergence Technologies (ACT) • Hyderabad

On-site
INR 4,000,000 - 8,000,000
Site Reliability Engineer (SRE) - Google Cloud Platform
Site Reliability Engineer (SRE) - Google Cloud Platform

Aziro • Hyderabad

Hybrid
INR 1,500,000 - 3,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000