Senior SRE

Nisum

Hyderabad

On-site

INR 2,800,000 - 4,200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nisum is seeking an experienced Senior Site Reliability Engineer (SRE) to own and optimize Kubernetes-based production systems on AWS. The role demands hands-on expertise in Kubernetes, Docker, CI/CD, monitoring, automation, and incident response.

You will manage AWS cloud infrastructure, design resilient architectures, automate operations with Python and Bash, and work with modern monitoring stacks like Prometheus, Grafana, Splunk, and CloudWatch.

Qualifications

  • 6+ years of SRE/production engineering experience.
  • Strong hands-on with Kubernetes, AWS, Docker, and CI/CD.
  • Experience in incident response and automation.
  • Exposure to AI-assisted SRE tooling preferred.

Responsibilities

  • Design, deploy, manage, and troubleshoot Kubernetes-based production.
  • Manage cloud infrastructure and services on AWS.
  • Work with Docker and Helm for containerization and deployments.
  • Develop and troubleshoot CI/CD pipelines using Jenkins.
  • Perform production troubleshooting across Linux, Kubernetes, apps, DBs, networking.
  • Automate tasks with Bash, Shell, and Python.
  • Monitor apps and infra with Splunk, Dynatrace, Grafana, Prometheus, ELK, CloudWatch.
  • Handle ITIL/ITSM processes for incidents and releases.
  • Use ServiceNow, Jira, and related tools for incident management.
  • Manage certificate renewals with Venafi / CertiS.
  • Support MySQL/SQL database troubleshooting.
  • Work with Akamai for traffic routing and issues.
  • Assist in production deployments, on-call support, root-cause analysis.
  • Identify automation opportunities to improve reliability.

Skills

Kubernetes
AWS
Docker & Helm
Jenkins / CI-CD
Linux (AWS Linux / Ubuntu)
Bash / Python
Monitoring tooling
ITIL / ITSM
Git (Bitbucket / GitHub / GitLab)
Apache Kafka
Akamai
Venafi / CertiS
MySQL / SQL
AI-assisted SRE

Tools

GitHub Copilot
Agentic AI

Job description

Senior SRE Kubernetes / AWS

Experience: 6+ Years

Location: Hyderabad

Employment Type: Contract-to-Hire (CTH) 1 Year

Job Overview

We are looking for an experienced Senior Site Reliability Engineer (SRE) with strong hands‑on expertise in Kubernetes, AWS, Docker, CI/CD, monitoring, automation, and production support.

The ideal candidate should have strong experience managing Kubernetes‑based production environments, troubleshooting critical incidents, improving system reliability, and automating operational processes. Exposure to AI‑powered SRE tools and Agentic AI for incident monitoring and remediation is required.

Key Responsibilities

  • Design, deploy, manage, and troubleshoot Kubernetes‑based applications and infrastructure in production.
  • Manage cloud infrastructure and services on AWS.
  • Work with Docker and Helm for containerization and Kubernetes deployments.
  • Develop, maintain, and troubleshoot CI/CD pipelines using Jenkins and Git‑based tools.
  • Perform production troubleshooting across Linux systems, Kubernetes, applications, databases, and networking.
  • Automate operational tasks using Bash/Shell scripting and Python.
  • Monitor applications and infrastructure using tools such as Splunk, Dynatrace, Grafana, Prometheus, ELK, and CloudWatch.
  • Handle Incident, Problem, Change, Release, Hot Fix, and ECRQ management following ITIL/ITSM processes.
  • Work with ServiceNow, Remedy, Jira, and/or Rally for service and incident management.
  • Manage certificate renewals using Venafi / CertiS.
  • Support database‑related troubleshooting involving MySQL and SQL.
  • Work with Akamai for traffic routing and related production issues.
  • Support messaging/event‑driven systems using Apache Kafka / Axon API.
  • Participate in production deployments, release activities, on‑call support, and root‑cause analysis.
  • Identify opportunities for automation and continuously improve system reliability, availability, and performance.

AI / SRE Automation Experience – Required

Candidates should have exposure to AI‑assisted SRE and automation, particularly:

  • GitHub Copilot for CI/CD pipeline scripting, Kubernetes YAML generation, cloud configuration, documentation, and automation.
  • Agentic AI for monitoring logs, identifying incidents, performing automated remediation, and escalating issues.
  • Exposure to AI‑powered fraud/risk platforms such as Visa Advanced Authorization, Mastercard Decision Intelligence, or Stripe Radar is preferred.

Mandatory Technical Skills

  • Kubernetes – Strong hands‑on production experience
  • AWS
  • Docker & Helm
  • Jenkins / CI-CD
  • Linux – AWS Linux / Ubuntu
  • Bash / Shell & Python
  • Monitoring – Splunk, Dynatrace, Prometheus, Grafana, ELK, CloudWatch
  • ITIL / ITSM – Incident, Problem, Change & Release Management
  • Git – Bitbucket / GitHub / GitLab
  • Apache Kafka
  • Akamai
  • Venafi / CertiS
  • MySQL / SQL
  • AI‑assisted SRE / Copilot / Agentic AI exposure
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE - AWS DevOPS Engineer
SRE - AWS DevOPS Engineer

Prowess Publishing • Hyderabad

On-site
INR 900,000 - 1,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hilabs • Bengaluru

On-site
INR 2,200,000 - 3,500,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC • Hyderabad, Bengaluru

Hybrid
INR 900,000 - 1,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mastercard • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
Senior SRE Engineer
Senior SRE Engineer

EPAM Systems • Gurugram District

On-site
INR 3,000,000 - 5,000,000