Remote Lead SRE - Cloud & Reliability

Worky

Eden Prairie (MN)

Hybrid

USD 113,000 - 193,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work

Job summary

UnitedHealth Group's OptumRx Digital is seeking a Lead Site Reliability Engineer to own the reliability, performance, and resilience of critical pharmacy services and cloud infrastructure. You will drive measurable improvements in availability and customer experience across high‑throughput digital platforms.

You will mentor engineers, define SRE best practices, and lead incident response, RCA, and post‑mortems, while partnering with product and software teams to optimize the digital supply

Qualifications

  • 7+ years of professional experience in Site Reliability Engineering, DevOps, or Software Engineering supporting enterprise applications and cloud infrastructure.
  • 5+ years of experience managing, monitoring, and maintaining production cloud applications and microservices (e.g., AWS, Azure, or GCP)
  • 4+ years of experience with containerized application environments and orchestration frameworks (e.g., Docker, Kubernetes)
  • 3+ years of experience setting up application performance monitoring (APM) and enterprise observability platforms (e.g., Datadog, Dynatrace, Prometheus, Grafana, or Splunk)

Responsibilities

  • Drive overall technical accountability for the availability, performance, and operational well-being of OptumRx Digital applications, microservices, and supporting infrastructure
  • Partner with software engineering and product leadership to influence digital supply chain technology workflows, optimizing transactional throughput, system integrations, and application health
  • Architect and implement robust application performance monitoring (APM), logging, and observability solutions to proactively detect, diagnose, and resolve application and service degradations
  • Automate cloud infrastructure, deployment pipelines, and operational processes using Infrastructure as Code (IaC) and modern CI/CD practices
  • Establish, track, and champion key application reliability metrics, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across applications and services
  • Lead end-to-end incident management, root cause analysis (RCA), and post-mortem actions to drive continuous improvement and eliminate recurring application failures across supply chain platforms
  • Provide technical leadership, mentorship, and guidance to engineering teams, fostering a culture of operational rigor, engineering quality, and continuous delivery

Skills

Site Reliability Engineering
DevOps
Cloud infrastructure
Observability
Leadership

Tools

Docker
Kubernetes
Datadog
Dynatrace
Prometheus
Grafana
Splunk
Terraform
CloudFormation
Ansible

Job description

UnitedHealth Group's OptumRx Digital is seeking a Lead Site Reliability Engineer to own the reliability, performance, and resilience of critical pharmacy services and cloud infrastructure. You will drive measurable improvements in availability and customer experience across high‑throughput digital platforms.

You will mentor engineers, define SRE best practices, and lead incident response, RCA, and post‑mortems, while partnering with product and software teams to optimize the digital supply

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability Engineer (Hybrid/Remote)
Senior SRE: Cloud Reliability Engineer (Hybrid/Remote)

Omnicell • Austin (TX)

Hybrid
USD 130,000 - 185,000
Remote or hybrid work
Up to 10% travel
Senior SRE Leader: AI-Powered Reliability for Cloud
Senior SRE Leader: AI-Powered Reliability for Cloud

Optum • Minnetonka (MN)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Incentive and recognition programs
401k contribution
Remote Cloud SRE: Build Resilient Health-Tech Infra
Remote Cloud SRE: Build Resilient Health-Tech Infra

UnitedHealth Group • Eden Prairie (MN)

Remote
Confidential
Remote work flexibility
Comprehensive benefits package
Equity stock purchase
+1
Site Reliability Engineer – Cloud Platforms (Hybrid)
Site Reliability Engineer – Cloud Platforms (Hybrid)

Omnicell • Cranberry Township

Hybrid
USD 120,000 - 180,000
Senior SRE & Software Engineer: Reliability Lead
Senior SRE & Software Engineer: Reliability Lead

CVS Health • Woonsocket (RI)

On-site
USD 93,000 - 204,000
Senior SRE & Software Engineer: Domain Reliability Lead
Senior SRE & Software Engineer: Domain Reliability Lead

Hispanic Alliance for Career Enhancement • Richardson (TX)

On-site
USD 93,000 - 204,000
Lead Site Reliability Engineer, Chief Digital Office
Lead Site Reliability Engineer, Chief Digital Office

Worky • Eden Prairie (MN)

Hybrid
USD 113,000 - 193,000
Remote work
Remote AWS SRE Lead: Cloud Reliability & AI Ops
Remote AWS SRE Lead: Cloud Reliability & AI Ops

JobCubby • Northern (KY)

Hybrid
USD 73,000 - 130,000
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
AI-Driven Principal SRE — Remote & Cloud Reliability
AI-Driven Principal SRE — Remote & Cloud Reliability

UnitedHealth Group • Eden Prairie (MN)

Hybrid
Confidential
Remote work options