Position Overview
The Pharmacy Benefit Services+ Technology organization seeks a Site Reliability Engineer (SRE) – Automation, Self‑Healing & AI/AIOps to join our team. This Band 4 Contributor role is a senior, hands‑on position responsible for driving enterprise reliability outcomes, reducing operational toil, and enabling scalable SRE adoption across both legacy platforms and modern cloud‑native systems.
Key Responsibilities
- Lead the design and implementation of intelligent, automated, and AI‑assisted reliability solutions that ensure systems are resilient, observable, self‑healing, and continuously improving.
- Build automation‑first and agentic SRE capabilities, including:
- Self‑healing workflows that automatically detect, diagnose, and remediate failure.
- AI‑driven operational intelligence (AIOps) for anomaly detection, alert correlation, incident triage, and guided remediation.
- Standardized SRE enablement platforms (SLO automation, reliability scorecards, FMEA workflows) that can be adopted at scale with minimal friction.
- Collaborate with application teams, platform engineering, DevOps, infrastructure, QE, and IT leadership to embed reliability into the SDLC and runtime operation.
- Improve system availability and resilience through proactive reliability engineering and automation.
- Reduce incidents and accelerate MTTR via self‑healing and AI‑assisted operational processes.
- Increase developer productivity by eliminating manual operational toil.
- Support faster, safer releases by integrating SRE controls into CI/CD pipelines.
- Deliver measurable reliability improvements, such as reductions in MTTD/MTTR, decreased incident frequency, improved SLO compliance, and healthier error‑budget consumption.
Qualifications
- Skills: Site Reliability Engineering: Deep hands‑on experience with SLOs, error budgets, incident management, and production operations.
- Automation & Software Engineering: Strong development skills in Python, Go, Java, or similar for production‑grade automation.
- Self‑Healing Systems: Proven experience designing and implementing automated remediation and closed‑loop recovery workflows.
- AI / AIOps: Experience applying AI/ML to operations, such as anomaly detection, alert correlation, predictive analysis, or intelligent remediation.
- Observability: Expertise with Dynatrace, Prometheus, Grafana, Splunk, AppDynamics, or equivalent platforms.
- Cloud & Distributed Systems: Understanding of AWS, Azure, GCP, microservices, and Kubernetes.
- CI/CD & DevOps: Experience integrating reliability checks and automation into delivery pipelines.
- Infrastructure as Code: Terraform, CloudFormation, or similar.
- Legacy + Modern Engineering: Ability to support and modernize reliability practices across monoliths, batch jobs, messaging, and mainframe‑integrated systems.
- Leadership & Influence: Ability to lead through influence, mentor others, and drive adoption across multiple teams.
Required Experience & Education
- Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent experience).
- 7+ years of experience in SRE, DevOps, platform engineering, or production software engineering roles.
- Demonstrated success delivering enterprise‑scale automation, self‑healing, and reliability improvements.
Desired Experience
- Experience building or contributing to enterprise SRE enablement platforms (SLO automation, reliability scorecards, FMEA workflows).
- Hands‑on experience with chaos engineering and resilience testing in production‑like environments.
- Familiarity with ServiceNow / CMDB / service modeling to support operational readiness and dependency visibility.
- Experience applying Generative AI for operational use cases such as runbook generation, incident summarization, and knowledge retrieval.
- Demonstrated delivery of quantifiable reliability improvements (e.g., MTTR reduction, incident volume reduction, improved SLO adherence).
- Experience mentoring engineers and shaping an automation‑first, reliability‑driven culture.
Location
Full‑time position, working 40 hours per week. Expected overlap with US hours as appropriate. Primarily based in the Innovation Hub in Hyderabad, India in a hybrid working model (3 days WFO).