Principal Site Reliability Engineer - Remote

Optum

Eden Prairie (MN)

On-site

USD 135,000 - 231,000

Full time

15 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote work within the U.S.
Office in Minneapolis or DC four days/

Job summary

Optum is seeking a Principal Site Reliability Engineer to evolve our reliability platform across Azure and AWS, combining modern SRE with AI-assisted operations. You will mentor engineers, define reliability standards, and drive next‑generation observability and automation across critical services.

The role offers remote work from anywhere in the U.S.; however, in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Qualifications

  • 10+ years of experience in software engineering, platform engineering, DevOps, or SRE roles.
  • 3+ years in a principal, staff, lead, or senior technical leadership role.
  • 5+ years of experience with cloud platforms and container orchestration (Azure or AWS).
  • 3+ years of experience with observability tools (OpenTelemetry, Prometheus, Grafana, Datadog).
  • 1+ years designing production automation or AI-assisted workflows for incident response.

Responsibilities

  • Build AI-assisted SRE capabilities that accelerate incident detection and recovery.
  • Connect observability data into actionable operational context.
  • Design human-in-the-loop workflows for safe mitigation and audits.
  • Standardize SLIs/SLOs and reliability scorecards across services.
  • Lead resiliency testing, DR exercises, and chaos engineering.
  • Mentor engineers and drive reliability improvements.

Skills

SRE leadership
Cloud & containers
Observability mastery
AI-assisted workflows
Incident response automation

Education

Bachelor's degree in CS/IT/Engineering

Tools

Azure
AWS
Kubernetes
OpenTelemetry
Prometheus
Grafana
Datadog
Terraform
Pulumi
Ansible
Helm

Job description

Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.

Are you passionate about reimagining operations through AI? Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform by combining modern SRE practices with AI-assisted operations. You’ll design systems that help engineers detect issues faster, automate response workflows, improve resiliency, and transform operational data into actionable insights. As a technical leader, you’ll influence reliability standards, mentor engineers, and drive next‑generation observability and automation strategies across Azure and AWS environments.

You’ll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges.

For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Primary Responsibilities
  • Build AI-assisted SRE capabilities that accelerate incident detection, triage, mitigation, and recovery
  • Connect observability, deployment, runbook, ownership, and incident data into actionable operational context
  • Design human-in-the-loop workflows for safe mitigation, approvals, recovery verification, and auditability
  • Standardize OpenTelemetry, SLIs, SLOs, error budgets, and reliability scorecards across critical services
  • Improve alert quality by reducing noise, clarifying customer impact, and identifying likely causes faster
  • Lead resiliency testing, DR exercises, chaos engineering, and automated recovery validation
  • Mentor engineers and lead cross-functional reliability improvements across Optum Financial

You’ll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.

Required Qualifications
  • 10+ years of experience in software engineering, platform engineering, DevOps, or SRE roles
  • 3+ years of experience in a principal, staff, lead, or senior technical leadership role
  • 5+ years of experience with cloud platforms and container orchestration, preferably Azure or AWS
  • 3+ years of experience with observability tools such as OpenTelemetry, Prometheus, Grafana, Datadog, or similar platforms
  • 1+ years of experience designing production automation, tooling, or AI-assisted workflows for incident response or operational decision-making
Preferred Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field
  • Experience with LLM-based systems, AI agents, RAG, tool orchestration, evaluations, or guardrails
  • Experience with resiliency engineering, disaster recovery, chaos engineering, or recovery validation
  • Experience with infrastructure as code and automation tools such as Terraform, Pulumi, Ansible, Helm, or Kubernetes operators
  • Solid background in incident command, runbooks, postmortems, production readiness, and reliability governance
  • All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy

Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you’ll find a far-reaching choice of benefits and incentives. The salary for this role will range from $134,600 - $230,800 annually based on full-time employment. We comply with all minimum wage laws as applicable.

At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.

UnitedHealth Group is an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.

UnitedHealth Group is a drug - free workplace. Candidates are required to pass a drug test before beginning employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer - Remote
Principal Site Reliability Engineer - Remote

UnitedHealth Group • Eden Prairie (MN)

Hybrid
Confidential
Remote work options
Site Reliability Engineer - Remote
Site Reliability Engineer - Remote

Optum • Eden Prairie (MN)

On-site
USD 73,000 - 130,000
Comprehensive benefits package
Equity stock purchase
401k contribution
+1
AWS Cloud Site Reliability Engineer
AWS Cloud Site Reliability Engineer

JobCubby • Northern (KY)

Hybrid
USD 73,000 - 130,000
Lead Site Reliability Engineer, Chief Digital Office
Lead Site Reliability Engineer, Chief Digital Office

Worky • Eden Prairie (MN)

Hybrid
USD 113,000 - 193,000
Remote work
Senior Software Engineer
Senior Software Engineer

UnitedHealth Group • Washington

Hybrid
Confidential
Comprehensive benefits package
Equity stock purchase
401k contribution
+1
Director, Technology Delivery and AI Engineering - Remote
Director, Technology Delivery and AI Engineering - Remote

UnitedHealth Group • Eden Prairie (MN)

Hybrid
Confidential
Comprehensive benefits
Equity stock purchase program
401k contribution
Lead Software Engineer - Remote
Lead Software Engineer - Remote

Optum • Eden Prairie (MN)

Hybrid
USD 113,000 - 193,000
Principal Software Engineer (AI Platform & Full Stack Engineering) - Hybrid in MN
Principal Software Engineer (AI Platform & Full Stack Engineering) - Hybrid in MN

Optum • Minnetonka (MN)

Hybrid
USD 134,000 - 231,000
Comprehensive benefits package
Equity stock purchase
401(k) contribution
Lead Software Engineer - Remote
Lead Software Engineer - Remote

UnitedHealth Group • Alpharetta (GA)

Hybrid
Confidential
Comprehensive benefits package
Equity stock purchase
401k contribution
Senior Data Analytics Engineer - Remote
Senior Data Analytics Engineer - Remote

Optum • Eden Prairie (MN)

Hybrid
USD 91,000 - 164,000
Comprehensive benefits package
401k contribution
Equity stock purchase