Senior Site Reliability Engineer

Optum

Chennai District

On-site

INR 3,000,000 - 6,000,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optum is hiring a Senior Site Reliability Engineer in Chennai, India to lead AI Ops initiatives and ensure reliable cloud infrastructure. You will design and implement AI-powered solutions, automate infrastructure with Terraform and GitHub Actions, and mentor engineers across cross-functional teams.

Strong Python and Kubernetes experience are essential, along with on-call flexibility. The role emphasizes scalable, secure, and responsible AI deployments, collaboration with development and SRE

Qualifications

  • Bachelor's degree in information systems, Computer Science, Engineering, or related field or equivalent certification.
  • 5+ years of software engineering and SRE experience in a public cloud environment (GCP, AWS, Azure).
  • 3+ years hands-on experience with Python and Terraform development.
  • 2+ years delivering AI/ML or Generative AI solutions in production.
  • 2+ years of experience managing Kubernetes environments (EKS, AKS, GKE, or self-hosted).
  • Available to work rotating 24x7 primary and secondary on-call shifts.

Responsibilities

  • Lead end-to-end design and implementation of AI Ops solutions from concept through production with an emphasis on responsible AI practices.
  • Contribute to the improvement of our SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environments.
  • Automate Infrastructure & Operations: Develop Infrastructure as Code using Terraform & GitHub Actions while adhering to best practices.
  • Define and own solution architecture for RAG pipelines, agentic workflows, tool calling, orchestration, and conversational context management.
  • Design, develop, and deploy AI-powered solutions to address complex business challenges across enterprise scale using RAG based solutions.
  • Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineers.
  • Design, code, test, and operate software using Python & Node.js.
  • Leverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous operational efficiency.
  • Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives.

Skills

Python
Terraform
Kubernetes
Cloud platforms
On-call readiness

Education

Bachelor's degree in information systems, Computer Science, Engineering, or related field

Tools

GitHub Actions
Terraform
Kubernetes

Job description

Improve the lives of others while Caring. Connecting. Growing together.

Job Description - Senior Site Reliability Engineer (2378612)

Senior Site Reliability Engineer - 2378612

Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.

Primary Responsibilities:
  • Lead end-to-end design and implementation of AI Ops solutions from concept through production with an emphasis on responsible AI practices
  • Contribute to the improvement of our SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environments
  • Automate Infrastructure & Operations: Develop Infrastructure as Code using Terraform & GitHub Actions while adhering to best practices
  • Define and own solution architecture for RAG pipelines, agentic workflows, tool calling, orchestration, and conversational context management
  • Design, develop, and deploy AI-powered solutions to address complex business challenges across enterprise scale using RAG based solutions
  • Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineers
  • Design, code, test, and operate software using Python & Node.js
  • Leverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous operational efficiency
  • Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
  • Bachelor’s degree in information systems, Computer Science, Engineering, or related field or equivalent certification
  • 5+ years of overall software engineering and Site Reliability Engineering (SRE) experience in a public cloud environment (GCP, AWS, Azure)
  • 3+ years of demonstrated hands‑on experience with Python and Terraform based development
  • 2+ years delivering AI/ML or Generative AI solutions in production
  • 2+ years of experience managing Kubernetes environments (EKS, AKS, GKE, or self-hosted)
  • Available to work rotating 24x7 primary and secondary on-call shifts
Preferred Qualifications:
  • Hands-on experience with cloud infrastructure automation, observability tools, and SRE best practices for AI workloads
  • Experience leading globally distributed technical teams
  • Experience with vector search and enterprise search solutions
  • CI/CD (GitHub Actions preferred) & Infrastructure as Code experience with Terraform
  • Proven experience building Retrieval Augmented Generation (RAG) pipelines, Agentic AI, or multi-step AI workflows
  • Proven experience building evaluation and monitoring frameworks for AI quality
  • Proven effective communication skills with ability to explain complex technical concepts to diverse stakeholders


At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone–of every race, gender, sexuality, age, location and income–deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes — an enterprise priority reflected in our mission.

UnitedHealth Group is committed to working with and providing reasonable accommodations to individuals with physical and mental disabilities. If you need special assistance or accommodation for any part of the application process, please call 1-866-566-8715 to be connected to Recruitment Services. Recruitment Services hours of operation are 7 a.m. to 7 p.m. CT, Monday through Friday.

UnitedHealth Group is a registered service mark of UnitedHealth Group, Inc. The UnitedHealth Group name with the dimensional logo, as well as the dimensional logo alone, are both service marks for the UnitedHealth Group, Inc.

Diversity creates a healthier atmosphere: UnitedHealth Group is an Equal Employment Opportunity/Affirmative Action employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, national origin, protected veteran status, disability status, sexual orientation, gender identity or expression, marital status, genetic information, or any other characteristic protected by law.

UnitedHealth Group is a drug-free workplace. Candidates are required to pass a drug test before beginning employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optum India • Chennai District

On-site
INR 350,000 - 750,000
Software Engineering Lead- Devops / SRE
Software Engineering Lead- Devops / SRE

Optum India • Chennai District

On-site
INR 2,200,000 - 3,600,000
Site Reliability Engineer
Site Reliability Engineer

Optum India • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior AI ML Engineer
Senior AI ML Engineer

Optum • Gurugram District

On-site
INR 4,200,000 - 7,000,000
Senior Software Engineering Lead
Senior Software Engineering Lead

UnitedHealth Group • India

On-site
Confidential
Senior Software Engineering Lead- Typescript, React, Terraform
Senior Software Engineering Lead- Typescript, React, Terraform

Optum India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software Engineering Lead - Python or Java
Software Engineering Lead - Python or Java

UnitedHealth Group • Chennai District

On-site
Confidential
Senior Software Engineer I - JAVA, React, SpringBoot, MicroService, Gen AI
Senior Software Engineer I - JAVA, React, SpringBoot, MicroService, Gen AI

Optum India • Hyderabad

On-site
INR 4,200,000 - 6,000,000
Software Engineering Lead - Python or Java
Software Engineering Lead - Python or Java

Optum India • Chennai District

On-site
INR 1,500,000 - 2,600,000
Senior Software Engineering Lead- Typescript, React, Terraform
Senior Software Engineering Lead- Typescript, React, Terraform

UnitedHealth Group • Bengaluru

On-site
Confidential