AI-Driven Principal SRE — Remote & Cloud Reliability

UnitedHealth Group

Eden Prairie (MN)

Hybrid

Confidential

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work options

Job summary

UnitedHealth Group's Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform. You will design AI-assisted workflows, improve observability, and drive automation across Azure and AWS environments.

You’ll mentor engineers, set reliability standards, and partner with cross-functional teams to accelerate incident detection, triage, recovery, and governance while maintaining security and compliance.

Qualifications

  • 10+ years of experience in software engineering, platform engineering, DevOps, or SRE roles.
  • 3+ years of experience in a principal, staff, lead, or senior technical leadership role.
  • 5+ years of experience with cloud platforms and container orchestration, preferably Azure or AWS.
  • 3+ years of experience with observability tools such as OpenTelemetry, Prometheus, Grafana, Datadog, or similar platforms.

Responsibilities

  • Build AI-assisted SRE capabilities to accelerate incident detection, triage, mitigation, and recovery.
  • Connect observability, deployment, runbook, ownership, and incident data into actionable context.
  • Design human-in-the-loop workflows for safe mitigation, approvals, recovery verification, and auditability.
  • Standardize OpenTelemetry, SLIs, SLOs, error budgets, and reliability scorecards across critical services.
  • Improve alert quality by reducing noise and clarifying customer impact.
  • Lead resiliency testing, DR exercises, chaos engineering, and automated recovery validation.
  • Mentor engineers and drive reliability improvements across Optum Financial.

Skills

Software engineering
Platform engineering
DevOps
SRE
Cloud (Azure/AWS)
Observability
AI-assisted automation
Leadership

Education

Bachelor's degree in CS/IT/Engineering

Tools

OpenTelemetry
Prometheus
Grafana
Datadog
Kubernetes
Terraform
Ansible

Job description

UnitedHealth Group's Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform. You will design AI-assisted workflows, improve observability, and drive automation across Azure and AWS environments.

You’ll mentor engineers, set reliability standards, and partner with cross-functional teams to accelerate incident detection, triage, recovery, and governance while maintaining security and compliance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal SRE: AI-Driven Reliability Leader (Remote)
Principal SRE: AI-Driven Reliability Leader (Remote)

Papa John'S International, Inc. • Kalispell (MT)

On-site
USD 150,000 - 190,000
Comprehensive benefits package
Equity stock purchase
401(k) contribution
+1
Remote Principal SRE — AI-Driven Reliability Leader
Remote Principal SRE — AI-Driven Reliability Leader

Optum • Eden Prairie (MN)

Remote
USD 135,000 - 231,000
Senior SRE Leader: AI-Powered Reliability for Cloud
Senior SRE Leader: AI-Powered Reliability for Cloud

Optum • Minnetonka (MN)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Incentive and recognition programs
401k contribution
SRE Principal: AI-Driven Reliability Leader
SRE Principal: AI-Driven Reliability Leader

UnitedHealth Group • Minnetonka (MN)

Remote
Confidential
Comprehensive benefits package
Incentive and recognition programs
Equity stock purchase
+1
Principal Site Reliability Engineer - Remote
Principal Site Reliability Engineer - Remote

UnitedHealth Group • Eden Prairie (MN)

Hybrid
Confidential
Remote work options
Remote SRE Lead — AI Reliability & Observability
Remote SRE Lead — AI Reliability & Observability

United States Digital Space LLC • United States

Remote
USD 198,000 - 303,000
AI-First Staff SRE: Reliability Architect & Incident Lead
AI-First Staff SRE: Reliability Architect & Incident Lead

EarnIn • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
AI-Driven Site Reliability Engineer – Remote
AI-Driven Site Reliability Engineer – Remote

Upstart • Austin (TX), San Francisco (CA), New York (NY)

Hybrid
USD 142,000 - 197,000
Competitive pay
Annual equity grants
401(k) matching
+2
Senior SRE Platform Engineer – AI‑Driven Reliability
Senior SRE Platform Engineer – AI‑Driven Reliability

Socket.dev • Denver (CO)

Hybrid
USD 140,000 - 190,000
Principal Site Reliability Engineer – Hybrid Multi‑Cloud
Principal Site Reliability Engineer – Hybrid Multi‑Cloud

PowerPlan, Inc. • Smyrna (GA)

Hybrid
USD 140,000 - 200,000
Hybrid work model
Onsite office in Smyrna, GA