SRE Technical Specialist

HCL Technologies Limited

Hyderabad

On-site

INR 4,000,000 - 7,000,000

Full time

25 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

HCL Technologies Limited is seeking an experienced Site Reliability Engineer (SRE) to manage production Kubernetes environments and ensure platform reliability in a Hyderabad-based role. The ideal candidate will bring 10+ years in SRE, with strong expertise in Kubernetes, Linux, cloud platforms (AWS/Azure/GCP), observability, automation, and production support.

You will lead 24x7 support, incident management, RCA, and SLA/SLO compliance, while automating tasks via Bash, Python, Terraform, and

Qualifications

  • Experience managing production Kubernetes environments and cluster operations.

Responsibilities

  • Manage and maintain Kubernetes clusters, deployments, upgrades, capacity planning, RBAC, networking, storage, Ingress, ConfigMaps, Secrets, Services, Persistent Volumes, StatefulSets, DaemonSets, Jobs/CronJobs, Helm, and Autoscaling (HPA/VPA).
  • Provide 24x7 production support, incident management, RCA, postmortems, service restoration, and SLA/SLO compliance.
  • Monitor and improve platform reliability using Prometheus, Grafana, Loki, Elastic Stack, OpenTelemetry, and AlertManager.
  • Troubleshoot Kubernetes, Linux, container (Docker/OCI), networking (DNS, Load Balancers, TLS, Ingress), cloud, and infrastructure issues.
  • Automate operational tasks through Bash, Python, Terraform, and Ansible, and support Infrastructure as Code practices.
  • Support CI/CD and release management using GitHub Actions, GitLab CI, Jenkins, and ArgoCD (preferred).
  • Perform patching, cluster maintenance, security updates, backups, disaster recovery validation, and platform upgrades.
  • Create runbooks, operational documentation, dashboards, alerts, and capacity planning reports.
  • Collaborate with Development, Platform Engineering, Security, Networking, Cloud Operations, and DevOps teams to improve system resilience and operational efficiency.

Skills

Kubernetes
Linux
Cloud Platforms
Observability
Automation
Production Support
Incident Management
CI/CD
GitOps
Terraform
Ansible
Python
Bash

Tools

Prometheus
Grafana
Loki
Elastic Stack
OpenTelemetry
AlertManager
GitHub Actions
GitLab CI
Jenkins
ArgoCD
Terraform
Ansible
Docker

Job description

Seeking an experienced Support Specialist having experience of 10+ years working as Site Reliability Engineer (SRE) with strong expertise in Kubernetes, Linux, Cloud Platforms (AWS/Azure/GCP), Observability, Automation, and Production Support. Responsible for managing and supporting production Kubernetes environments, ensuring platform reliability, availability, security, scalability, and operational excellence.

Key Responsibilities:

  • Manage and maintain Kubernetes clusters, including deployments, upgrades, capacity planning, RBAC, networking, storage, Ingress, ConfigMaps, Secrets, Services, Persistent Volumes, StatefulSets, DaemonSets, Jobs/CronJobs, Helm, and Autoscaling (HPA/VPA).
  • Provide 24x7 production support, incident management, RCA, postmortems, service restoration, and SLA/SLO compliance.
  • Monitor and improve platform reliability using Prometheus, Grafana, Loki, Elastic Stack, OpenTelemetry, and AlertManager.
  • Troubleshoot Kubernetes, Linux, container (Docker/OCI), networking (DNS, Load Balancers, TLS, Ingress), cloud, and infrastructure issues.
  • Automate operational tasks through Bash, Python, Terraform, and Ansible, and support Infrastructure as Code practices.
  • Support CI/CD and release management using GitHub Actions, GitLab CI, Jenkins, and ArgoCD (preferred).
  • Perform patching, cluster maintenance, security updates, backups, disaster recovery validation, and platform upgrades.
  • Create runbooks, operational documentation, dashboards, alerts, and capacity planning reports.
  • Collaborate with Development, Platform Engineering, Security, Networking, Cloud Operations, and DevOps teams to improve system resilience and operational efficiency.

Required Skills: Kubernetes Administration, Linux, Docker/OCI, AWS/Azure/GCP, Networking, CI/CD, GitOps, Observability, Incident Management, RCA, Automation, Terraform, Ansible, Bash, Python.

Preferred: CKA/CKS certification, Cloud certifications, Multi-cluster/Multi-region Kubernetes, Service Mesh (Istio/Linkerd), High Availability, Disaster Recovery, Security Hardening, Capacity Planning, Performance Tuning, Cost Optimization, Chaos Engineering, AI-assisted Observability.

Key Competencies: Strong troubleshooting, ownership, production support, customer focus, communication, collaboration, continuous improvement, and ability to perform under pressure.

Success Metrics: High platform availability, improved MTTR, reduced incidents and alert noise, SLA/SLO compliance, increased automation coverage, successful upgrades/maintenance, and customer satisfaction.

Other Requirements

Preferred Qualifications

  • Certified Kubernetes Security Specialist (CKS)
  • Cloud certifications (AWS/Azure/GCP)
  • Experience supporting multi-cluster Kubernetes environments.
  • Experience with service mesh technologies (Istio/Linkerd).

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE
Senior SRE

Nisum • Hyderabad

On-site
INR 2,800,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer (SRE) – Kubernetes / OpenShift
Site Reliability Engineer (SRE) – Kubernetes / OpenShift

Altiquence Communications • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Technical Support Engineer/SRE
Technical Support Engineer/SRE

Boldtek • India

On-site
INR 900,000 - 1,500,000
Hybrid work model
Growth opportunities
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000
Lead SRE
Lead SRE

UST • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
SRE - AWS DevOPS Engineer
SRE - AWS DevOPS Engineer

Prowess Publishing • Hyderabad

On-site
INR 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Smart Ims • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000