Infrastructure Associate Advisor

Evernorth Health Services

Hyderabad

Hybrid

INR 1,000,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evernorth Health Services is seeking a Site Reliability Engineer (SRE) for their team in Hyderabad, India. This senior role focuses on automation, self-healing, and AI/AIOps to enhance reliability across systems.

Candidates should have 7+ years in SRE or DevOps, strong automation skills, and familiarity with cloud systems. The position offers a hybrid work model, with expected overlap with US hours.

Join us to drive enterprise reliability outcomes and reduce operational toil.

Qualifications

  • 7+ years of experience in SRE, DevOps, platform engineering, or production software engineering roles.
  • Demonstrated success delivering enterprise-scale automation, self-healing, and reliability improvements.

Responsibilities

  • Lead the design and implementation of intelligent, automated, and AI-assisted reliability solutions.
  • Build automation-first and agentic SRE capabilities.
  • Collaborate with application teams and IT leadership to embed reliability.
  • Support faster, safer releases by integrating SRE controls into CI/CD pipelines.

Skills

Site Reliability Engineering
Automation & Software Engineering
Self-Healing Systems
AI / AIOps
Observability
Cloud & Distributed Systems
CI/CD & DevOps
Infrastructure as Code
Legacy + Modern Engineering
Leadership & Influence

Education

Bachelor’s degree in Computer Science, Engineering, or related field

Tools

Python
Go
Java
AWS
Azure
GCP
Kubernetes
Terraform
CloudFormation
Dynatrace
Prometheus
Grafana
Splunk
AppDynamics

Job description

Position Overview

The Pharmacy Benefit Services+ Technology organization seeks a Site Reliability Engineer (SRE) – Automation, Self‑Healing & AI/AIOps to join our team. This Band 4 Contributor role is a senior, hands‑on position responsible for driving enterprise reliability outcomes, reducing operational toil, and enabling scalable SRE adoption across both legacy platforms and modern cloud‑native systems.

Key Responsibilities
  • Lead the design and implementation of intelligent, automated, and AI‑assisted reliability solutions that ensure systems are resilient, observable, self‑healing, and continuously improving.
  • Build automation‑first and agentic SRE capabilities, including:
  • Self‑healing workflows that automatically detect, diagnose, and remediate failure.
  • AI‑driven operational intelligence (AIOps) for anomaly detection, alert correlation, incident triage, and guided remediation.
  • Standardized SRE enablement platforms (SLO automation, reliability scorecards, FMEA workflows) that can be adopted at scale with minimal friction.
  • Collaborate with application teams, platform engineering, DevOps, infrastructure, QE, and IT leadership to embed reliability into the SDLC and runtime operation.
  • Improve system availability and resilience through proactive reliability engineering and automation.
  • Reduce incidents and accelerate MTTR via self‑healing and AI‑assisted operational processes.
  • Increase developer productivity by eliminating manual operational toil.
  • Support faster, safer releases by integrating SRE controls into CI/CD pipelines.
  • Deliver measurable reliability improvements, such as reductions in MTTD/MTTR, decreased incident frequency, improved SLO compliance, and healthier error‑budget consumption.
Qualifications
  • Skills: Site Reliability Engineering: Deep hands‑on experience with SLOs, error budgets, incident management, and production operations.
  • Automation & Software Engineering: Strong development skills in Python, Go, Java, or similar for production‑grade automation.
  • Self‑Healing Systems: Proven experience designing and implementing automated remediation and closed‑loop recovery workflows.
  • AI / AIOps: Experience applying AI/ML to operations, such as anomaly detection, alert correlation, predictive analysis, or intelligent remediation.
  • Observability: Expertise with Dynatrace, Prometheus, Grafana, Splunk, AppDynamics, or equivalent platforms.
  • Cloud & Distributed Systems: Understanding of AWS, Azure, GCP, microservices, and Kubernetes.
  • CI/CD & DevOps: Experience integrating reliability checks and automation into delivery pipelines.
  • Infrastructure as Code: Terraform, CloudFormation, or similar.
  • Legacy + Modern Engineering: Ability to support and modernize reliability practices across monoliths, batch jobs, messaging, and mainframe‑integrated systems.
  • Leadership & Influence: Ability to lead through influence, mentor others, and drive adoption across multiple teams.
Required Experience & Education
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent experience).
  • 7+ years of experience in SRE, DevOps, platform engineering, or production software engineering roles.
  • Demonstrated success delivering enterprise‑scale automation, self‑healing, and reliability improvements.
Desired Experience
  • Experience building or contributing to enterprise SRE enablement platforms (SLO automation, reliability scorecards, FMEA workflows).
  • Hands‑on experience with chaos engineering and resilience testing in production‑like environments.
  • Familiarity with ServiceNow / CMDB / service modeling to support operational readiness and dependency visibility.
  • Experience applying Generative AI for operational use cases such as runbook generation, incident summarization, and knowledge retrieval.
  • Demonstrated delivery of quantifiable reliability improvements (e.g., MTTR reduction, incident volume reduction, improved SLO adherence).
  • Experience mentoring engineers and shaping an automation‑first, reliability‑driven culture.
Location

Full‑time position, working 40 hours per week. Expected overlap with US hours as appropriate. Primarily based in the Innovation Hub in Hyderabad, India in a hybrid working model (3 days WFO).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Arch Systems • Hyderabad

On-site
INR 2,800,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SourcingXPress • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Site Reliability Engineer II
Site Reliability Engineer II

United States Digital Space LLC • Karnataka

On-site
INR 800,000 - 1,200,000
Associate Site Reliability Engineer II
Associate Site Reliability Engineer II

MetLife • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Resilience and Reliability Engineer
Resilience and Reliability Engineer

EY • Pune District, Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC • Hyderabad, Bengaluru

Hybrid
INR 900,000 - 1,400,000
Senior Manager – Site Reliability Engineering (SRE)- Remote
Senior Manager – Site Reliability Engineering (SRE)- Remote

First Advantage • Bengaluru

On-site
INR 2,500,000 - 3,500,000
Comprehensive employee Leave policy
Career Development programs
Medical Insurance coverage
+2
Site Reliability Specialist - High Availability
Site Reliability Specialist - High Availability

Freelanceshop • Gwalior District

Hybrid
INR 1,200,000 - 2,000,000
Competitive salary
Health and life insurance
Flexible work arrangements
+2