Principal SRE Engineer

Eli Lilly and Company

Indianapolis (IN)

On-site

USD 66,000 - 158,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Lilly is seeking a Site Reliability Engineer in Indianapolis to build self-healing automation and durable fixes across a multi-application estate. You will encode SLOs, error budgets, and runbooks into automated processes and work closely with the Senior Principal SRE Engineer to decide when automation or manual steps are warranted.

This hands-on role focuses on execution, validation, and continuous improvement of resiliency patterns, incident response, and structured postmortems, ensuring

Qualifications

  • Bachelor's degree in Computer Science or a related technical engineering discipline.
  • 5+ years of engineering experience in Site Reliability/Production engineering.
  • Experience with self-healing automation or runbook-driven remediation in a multi-application estate.

Responsibilities

  • Design and implement self-healing automation and resilience patterns.
  • Perform chaos testing to validate automated remedies before production use.
  • Author and maintain remediation runbooks with safe execution and rollback steps.
  • Collaborate with Operations and automation teams to validate outcomes.
  • Lead or contribute to incident RCA and blameless post-mortems to drive durable fixes.

Skills

Self-healing automation
Chaos testing
AWS reliability services
Runbooks
Kubernetes
Observability

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
AWS
Splunk/Datadog/New Relic

Job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.


About The Technology Organization

Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable.


About The Team

Technology at Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business.


Role Summary

You build the self-healing automation, author the runbooks, and turn root-cause analysis into durable engineering fixes that let the production estate heal itself instead of paging a human. You work within the standards the Senior Principal SRE Engineer sets — SLOs, error budgets, observability — and you're the one who encodes them into working automation and documented procedure.


This is a hands-on individual-contributor role focused on execution and codification rather than cross-estate reliability strategy. You decide, in partnership with the Senior Principal SRE Engineer, which recurring patterns warrant a self-healing investment versus a documented manual runbook, and you build whichever is right.


You are an individual contributor. You do not manage people. You partner daily with the Senior Principal SRE Engineer, the agentic automation engineering team, and Operations on validating outcomes. Success is measured by self-healing coverage, runbooks authored and adopted, reduced recurrence of known failure modes, and the safety record of every automation you sign off.


What You'll Be Doing


  • Self-healing automation & resilience patterns

  • Design and build self-healing automation — circuit breakers, graceful degradation, automated remediation — for the failure modes that recur most across the estate.

  • Run resilience or chaos testing to validate that self-healing patterns behave correctly before they're trusted in production.

  • Continuously expand self-healing coverage as new failure modes are identified and proven safe to automate.

  • Partner with the Senior Principal SRE Engineer on which failure modes justify self-healing investment versus a documented manual runbook.

  • Runbook authorship & validation

  • Author and validate the remediation runbooks for the production estate: safe execution order, rollback steps, and exception handling for every documented fix.

  • Keep the runbook library current as systems, dependencies, and failure modes evolve, retiring runbooks that no longer apply.

  • Define and apply the graduation criteria that let a runbook move from human-executed to agent-assisted to autonomous.

  • RCA to durable fix

  • Lead or contribute to root-cause analysis for significant incidents, and drive the blameless postmortem process to a durable engineering fix — not just a narrative.

  • Convert recurring incident patterns into codified runbooks and, where appropriate, self-healing automation.

  • Track fix effectiveness against recurrence, and elevate to the Senior Principal SRE Engineer when a fix needs broader engineering investment.

  • Participate in high-severity incident response, including acting as incident commander for escalations within your area.

  • Partnership with agentic automation & operations

  • Partner with the Agentic Automation Engineering team on which fixes are safe to hand off as agent-assisted remediations, and on the confidence thresholds and human-in-the-loop boundaries that keep them safe.

  • Sign off on agent graduation criteria (accuracy over volume, zero P1/P2 caused) before an automation moves to a higher autonomy tier.

  • Partner with Operations on outcome validation, feeding what's learned back into the runbook library and self-healing patterns.

  • Incident response & regulated-environment practice

  • Ensure runbooks and self-healing automation meet Lilly's change-control, audit, and validated-environment standards.

  • Document procedures so that audit evidence falls out of normal operation, not a special exercise.

  • Mentor other reliability and automation engineers on runbook quality and self-healing design.

  • Contribute proven patterns back to the broader reliability practice, in partnership with the Senior Principal SRE Engineer and Senior Architect.


How You Will Succeed


  • Be recognized as the engineer who turns incidents into durable fixes, not repeat pages.

  • Demonstrate measurable growth in self-healing coverage and runbook adoption, with falling recurrence of known failure modes.

  • Maintain a clean safety record: automations you sign off don't cause P1/P2 incidents.

  • Build runbooks and automation that make good practice the default, not a personal habit.


Your Basic Qualifications


  • Bachelor's degree in Computer Science, Information Technology, or a related technical engineering discipline, including Software Engineering, Computer Engineering, Information Systems, Cybersecurity, Information Science, Network Engineering, Systems Engineering, Computer Information Systems (CIS), Management Information Systems (MIS), Cloud Computing, Data Science

  • 5+ years of progressive engineering experience, with meaningful time as a Site Reliability Engineer, Production Engineer, or equivalent, including hands-on ownership of self-healing automation or runbook-driven remediation for a multi-application production estate.

  • Production reliability experience in a regulated or audited environment (GxP, SOX, HIPAA, PCI, or equivalent), including familiarity with change-control discipline, audit evidence, and validated-system constraints.

  • Hands-on experience authoring and validating runbooks: safe execution order, rollback steps, and exception handling for real remediation procedures.

  • Demonstrated hands-on experience designing, implementing, and operating enterprise-scale SRE platforms, including observability solutions (Splunk, Datadog, New Relic, or Grafana/Prometheus), infrastructure-as-code with Terraform, CI/CD pipeline hardening, Kubernetes-based container platforms, and production workloads hosted on AWS, Azure, or GCP.

  • Experience designing self-healing patterns (circuit breakers, graceful degradation, automated remediation) and validating them before they're trusted in production.

  • Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT,J-1 , H-1B, TN, O-1, E-3, H-1B1, or L-1.


What You Should Bring


  • Hands-on experience designing self-healing automation and running chaos engineering or resilience-testing programs (AWS Fault Injection Service, Gremlin, LitmusChaos, or equivalent) tied to measurable reliability gains.

  • Deep AWS fluency across reliability-relevant services (EKS, ECS, Lambda, CloudWatch, X-Ray, Systems Manager, Route 53), and familiarity with AWS Well-Architected Reliability Pillar.

  • Experience with AIOps or agent-assisted operations, including designing the guardrails, confidence thresholds, and human-in-the-loop boundaries that make automated remediation safe in production.

  • Experience defining or signing off on autonomy/graduation criteria for automation moving toward self-sufficiency.

  • Prior experience building or scaling a runbook library across multiple applications or teams.

  • Experience operating in pharma, healthcare, financial services, or other regulated industries.

  • Track record of contributing meaningfully to root-cause analysis and blameless post-mortems, converting findings into durable engineering fixes measured by reduced recurrence rather than narrative quality.


Leadership Expectations


  • Treats every RCA as unfinished until it produces a durable fix, not just a narrative.

  • Combines engineering rigor with operational pragmatism in deciding what gets automated and what stays manual.

  • Leads through what they build and how they write, not through org-chart authority.

  • Comfortable telling a team that a fix isn't ready for autonomy yet, and showing them what's missing.

  • Treats runbook quality and knowledge-sharing as a first-class outcome, not an afterthought.


Additional Information

Availability to work flexible work hours is/may be required. This team supports continuous operations and may require non-standard work hours, including some work on weekends and holidays.


Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.


Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.


Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).


Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is


$66,000 - $158,400


Full-time equivalent employees also

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Lilly Group • Indianapolis (IN)

On-site
USD 66,000 - 158,000
401(k)
Pension
Vacation benefits
+4
Senior Principal SRE Engineering
Senior Principal SRE Engineering

Initial Therapeutics, Inc. • Indianapolis (IN)

On-site
USD 129,000 - 231,000
401(k)
Pension
Vacation benefits
+2
Principal SRE Engineer
Principal SRE Engineer

Initial Therapeutics, Inc. • Indianapolis (IN)

On-site
USD 150,000 - 190,000
Senior Principal SRE Engineering
Senior Principal SRE Engineering

BioSpace • Indianapolis (IN)

On-site
USD 129,000 - 231,000
401(k) plan
Health benefits
Senior Principal SRE Engineering
Senior Principal SRE Engineering

Eli Lilly and Company • Indianapolis (IN)

On-site
USD 129,000 - 231,000
401(k) plan
Health benefits
Well-being benefits
Senior Principal Agentic Engineer
Senior Principal Agentic Engineer

BioSpace • Indianapolis (IN)

On-site
USD 129,000 - 231,000
Associate Director- Automation Portfolio and Business Enablement
Associate Director- Automation Portfolio and Business Enablement

Eli Lilly and Company • Indianapolis (IN)

On-site
USD 129,000 - 189,000
Company bonus
401(k) plan
Pension
+2
Senior Principal Agentic Engineer
Senior Principal Agentic Engineer

Eli Lilly and Company • Indianapolis (IN)

On-site
USD 129,000 - 231,000
Associate Director- Automation Portfolio and Business Enablement
Associate Director- Automation Portfolio and Business Enablement

BioSpace • Indianapolis (IN)

On-site
USD 129,000 - 189,000
Engineer – Reliability
Engineer – Reliability

Tribonet • Indiana (PA), Northern (KY)

Hybrid
USD 66,000 - 171,600
Comprehensive benefits program
401(k) participation
Company bonus eligibility