Senior Site Reliability Engineer

Optum

Schaumburg (IL)

Remote

USD 92,000 - 164,000

Full time

3 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Optum is seeking a Senior Site Reliability Engineer to lead SRE practices, design scalable cloud-native platforms across AWS and Azure, and drive observability initiatives. You will develop alerts, dashboards, and reliable automation to reduce toil and improve system stability.

The role supports Kubernetes-based environments and CI/CD pipelines, with on-call rotation and collaboration across engineering, security, and platform teams. Remote work within the United States is possible.

Qualifications

  • Bachelor’s degree in a related field.
  • 7+ years in SRE/DevOps/Platform Engineering.
  • 4+ years cloud infra exposure (AWS/Azure/GCP).
  • 3+ years Kubernetes in production (EKS/AKS/GKE).
  • 3+ years IaC (Terraform or equivalent).
  • 3+ years observability/monitoring tools (Datadog/Splunk/Grafana/OpenTelemetry/Prometheus).
  • 3+ years CI/CD using GitHub Actions/Azure DevOps/Jenkins/ArgoCD or equivalent.
  • 3+ years scripting with Python/Bash/PowerShell.

Responsibilities

  • Lead SRE practices with SLIs/SLOs and error budgets to boost reliability.
  • Design and support cloud-native platforms across AWS, Azure, Kubernetes.
  • Build observability dashboards and alerting for end-to-end health visibility.
  • Lead incident response, RCA, and post-incident remediation.

Skills

SRE principles
Cloud platforms
Kubernetes
Observability
Terraform (IaC)
CI/CD pipelines
Python scripting
On-call experience

Education

Bachelor’s degree in Computer Science/Engineering/IT

Tools

Datadog
Splunk
Grafana
OpenTelemetry
Prometheus
Terraform
GitHub Actions
ArgoCD
Azure DevOps

Job description

Improve the lives of others while Caring. Connecting. Growing together.

Job Description - Senior Site Reliability Engineer (2385714)

Senior Site Reliability Engineer - 2385714

Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.

You’ll enjoy the flexibility to telecommute* from anywhere within the U.S. as you take on some tough challenges.

Primary Responsibilities:

  • Lead the implementation and continuous improvement of Site Reliability Engineering (SRE) practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to improve system reliability and operational excellence
  • Design, deploy, and support highly available, scalable, and resilient cloud-native platforms across AWS, Azure, and Kubernetes environments
  • Build and maintain observability solutions utilizing Datadog, Splunk, Grafana, OpenTelemetry, Prometheus, and related monitoring technologies to provide end-to-end visibility into platform and application health
  • Develop operational dashboards, intelligent alerting, service health scorecards, and reliability metrics to improve incident detection and operational awareness
  • Lead production incident response activities, root cause analysis (RCA), and post-incident remediation efforts to improve service stability and reduce recurrence
  • Develop and maintain Infrastructure as Code (IaC) solutions using Terraform and cloud automation technologies
  • Build automation and self-healing capabilities to reduce operational toil, improve platform resilience, and accelerate incident resolution
  • Partner with software engineering, architecture, security, and platform teams to improve production readiness, reliability, scalability, and performance
  • Support and optimize Kubernetes-based container platforms and microservices running in production environments
  • Design and implement CI/CD and GitOps practices utilizing GitHub Actions, ArgoCD, Azure DevOps, or equivalent deployment technologies
  • Drive adoption of AI-enabled operational capabilities including anomaly detection, intelligent alerting, incident automation, and operational analytics
  • Mentor engineers on SRE principles, observability, automation, operational excellence, and cloud-native best practices

You’ll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.

Required Qualifications:

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or related field
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or Software Engineering
  • 4+ years of hands‑on experience supporting cloud infrastructure in AWS, Azure, or GCP environments
  • 3+ years of hands‑on experience managing Kubernetes platforms including EKS, AKS, or GKE in production environments
  • 3+ years of Infrastructure as Code (IaC) experience using Terraform or equivalent automation technologies
  • 3+ years of hands‑on experience with observability and monitoring platforms such as Datadog, Splunk, Dynatrace, Grafana, Prometheus, OpenTelemetry, or similar solutions
  • 3+ years of experience implementing and supporting monitoring, logging, distributed tracing, alerting, SLIs, SLOs, and Error Budget frameworks
  • 3+ years of experience building and supporting CI/CD pipelines using GitHub Actions, Azure DevOps, Jenkins, ArgoCD, or equivalent technologies
  • 3+ years of experience with scripting and automation skills using Python, Bash, PowerShell, or similar languages
  • Ability to participate in rotating on‑call support schedules

Preferred Qualifications:

  • Experience supporting mission‑critical production systems and leading incident response and root cause analysis activities
  • Strong understanding of distributed systems, cloud‑native architectures, networking, security, IAM, encryption, and reliability engineering principles
  • Proven ability to collaborate effectively across engineering, platform, architecture, and security teams
  • Experience implementing enterprise observability solutions using Datadog APM, Splunk Observability Cloud, Dynatrace, Grafana, or OpenTelemetry
  • Experience with AIOps, intelligent alerting, anomaly detection, operational automation, and predictive analytics platforms
  • Experience supporting AI/ML, Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), or data‑intensive workloads in production environments
  • Experience with GitOps frameworks such as ArgoCD or Flux
  • Experience supporting multi‑region and multi‑cluster cloud deployments
  • Experience mentoring engineers and leading reliability improvements across multiple teams
  • Experience working within regulated environments such as Healthcare, HIPAA, SOC2, NIST, or FedRAMP
  • Industry certifications such as Certified Kubernetes Administrator (CKA), AWS Solutions Architect, Azure Solutions Architect Expert, HashiCorp Terraform Associate, or equivalent cloud certifications

*All employees working remotely will be required to adhere to UnitedHealth Group’s Telecommuter Policy

Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you’ll find a far-reaching choice of benefits and incentives. The salary for this role will range from $91,700 to $163,700 annually based on full‑time employment. We comply with all minimum wage laws as applicable.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.

At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone–of every race, gender, sexuality, age, location, and income–deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups, and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes — an enterprise priority reflected in our mission.

Diversity creates a healthier atmosphere: UnitedHealth Group is an Equal Employment Opportunity/Affirmative Action employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, national origin, protected veteran status, disability status, sexual orientation, gender identity or expression, marital status, genetic information, or any other characteristic protected by law.

UnitedHealth Group is a drug‑free workplace. Candidates are required to pass a drug test before beginning employment.

UnitedHealth Group is committed to working with and providing reasonable accommodations to individuals with physical and mental disabilities. If you need special assistance or accommodation for any part of the application process, please call 1-866-566-8715 to be connected to Recruitment Services. Recruitment Services hours of operation are 7 a.m. to 7 p.m. CT, Monday through Friday.

UnitedHealth Group is a registered service mark of UnitedHealth Group, Inc. The UnitedHealth Group name with the dimensional logo, as well as the dimensional logo alone, are both service marks for the UnitedHealth Group, Inc.

Diversity creates a healthier atmosphere: UnitedHealth Group is an Equal Employment Opportunity/Affirmative Action employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, national origin, protected veteran status, disability status, sexual orientation, gender identity or expression, marital status, genetic information, or any other characteristic protected by law.

UnitedHealth Group is a drug‑free workplace. Candidates are required to pass a drug test before beginning employment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

UnitedHealth Group • Schaumburg (IL)

Remote
USD 92,000 - 164,000
Telecommute within US
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Optum • Eden Prairie (MN)

On-site
USD 135,000 - 231,000
Comprehensive benefits
Equity stock purchase
401(k) contribution
Senior Manager Software Engineering - Remote
Senior Manager Software Engineering - Remote

Optum • Eden Prairie (MN)

Remote
USD 113,000 - 193,000
Comprehensive benefits package
Incentive and recognition programs
Equity stock purchase
+1
Sr Platform DevOps Engr, SRE - Remote
Sr Platform DevOps Engr, SRE - Remote

UnitedHealth Group • Eden Prairie (MN)

Hybrid
USD 92,000 - 164,000
Comprehensive benefits package
Equity stock purchase plan
401(k) contribution
Senior Manager, Software Engineering - Remote
Senior Manager, Software Engineering - Remote

Optum • Montgomery (AL)

Remote
USD 113,000 - 193,000
Site Reliability Engineer - Remote
Site Reliability Engineer - Remote

Optum • Eden Prairie (MN)

On-site
USD 73,000 - 130,000
Comprehensive benefits package
Equity stock purchase
401k contribution
+1
Senior Full Stack Software Engineer
Senior Full Stack Software Engineer

Optum • Eden Prairie (MN)

On-site
USD 92,000 - 164,000
Senior Software Engineer
Senior Software Engineer

Optum • Town of Wausau (WI)

On-site
USD 92,000 - 164,000
Comprehensive benefits
Equity stock purchase
401(k) contribution
+1
Principal Site Reliability Engineer
Principal Site Reliability Engineer

UnitedHealth Group • Eden Prairie (MN)

On-site
USD 135,000 - 231,000
Comprehensive benefits package
Equity stock purchase
401k contribution
Software Engineer
Software Engineer

Optum • La Crosse (WI)

On-site
USD 73,000 - 130,000
Comprehensive benefits
Equity stock purchase
401(k) contributions
+1