AWS Cloud Site Reliability Engineer

JobCubby

Northern (KY)

Hybrid

USD 73,000 - 130,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optum is seeking a Senior SRE to lead a team of engineers in designing and maintaining scalable cloud systems. You will implement monitoring strategies, drive automation, and collaborate on architecture and incident response to ensure high availability and performance.

The role requires 3+ years in AWS, Java/Spring, Python, and distributed data services, with a focus on reliability and IaC. Remote work is supported nationwide within the U.S., with in‑office requirements in some areas.

Qualifications

  • Bachelor’s degree in CS or IT or related field.
  • 3+ years with AWS Cloud services and Java/Spring.
  • 3+ years with distributed data services (DynamoDB/Athena).
  • 3+ years with AWS Cloud (S3, CloudWatch, ECS, Lambda, RDS, EMR).
  • 3+ years with CI/CD using GitHub Actions or similar.

Responsibilities

  • Lead and mentor a team of SREs to ensure high‑quality delivery and growth.
  • Design, build, and maintain scalable cloud systems with cloud‑native tech.
  • Develop monitoring, alerting, and observability strategies.
  • Automate tasks and drive IaC adoption.
  • Identify and resolve reliability risks, bottlenecks, and performance issues.
  • Collaborate with engineering and product on architecture and incident response.
  • Lead post‑incident reviews, root cause analysis, and continuous improvement.

Skills

AWS Java Spring
Scala
Python
DynamoDB/Athena
GitHub Actions
Cloud SDKs
Elastic APM
Kubernetes

Education

Bachelor’s degree in CS or IT related field

Tools

Kubernetes
Unix
Hadoop/HBase/Hive

Job description

Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by diversity and inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health equity on a global scale. Join us to start Caring. Connecting. Growing together.

You will be part of a world class identity matching solution building a state-of-the-art applications that is at the center of identity management for Optum Technology. You will have a true opportunity to change the healthcare landscape for the better. Role requires to provide 24×7 operational support to all production practices on holidays and weekends. Coordinate with various teams and raise support ticket for all issues, analyze root cause and assist in efficient resolution of all production processes. Maintain logs of all issues and ensure resolutions according to quality assurance tests for all production processes. Need to have good understanding of business processes within various systems used within the application. You will need to be ambitious and willing to work out of your comfort zone.

You’ll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office for a minimum of four days per week.

Primary Responsibilities
  • Lead and mentor a team of SREs to ensure high-quality delivery and professional growth
  • Design, build, and maintain scalable and reliable systems using cloud-native technologies
  • Develop and implement monitoring, alerting, and observability strategies to ensure optimal system performance and user experience
  • Automate operational tasks and drive infrastructure-as-code (IaC) adoption
  • Proactively identify and resolve reliability risks, bottlenecks, and performance issues
  • Leveraging AI
  • Collaborate with engineering and product teams on architecture, code reviews, and incident response
  • Lead post-incident reviews (blameless retrospectives), root cause analysis, and continuous improvement initiatives
  • Streamline migration processes, ensure consistency and enhance efficiency through automation, AI and innovative solutions
  • Define SLOs/SLIs, track error budgets, and report on system health to stakeholders
  • Ensure compliance and security standards are integrated into system operations
  • Stay current with emerging technologies and SRE best practices
  • Leverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous improvement
Required Qualifications
  • Bachelor’s degree OR CS OR IT related field
  • 3+ years of experience with Cloud SDKs with AWS using Java (spring boot microservices), Scala, and Python
  • 3+ years of experience with Distributed Data services (DynamoDB/Athena or similar)
  • 3+ years of experience with AWS Cloud: S3, CloudWatch, ECS, Lambda, RDS, EMR, AWS ECS
  • 3+ years of experience with CI/CD using GitHub Actions or similar
Preferred Qualifications
  • Experience in Unix, Hadoop, HBase and Hive
  • Experience working with offshore and onsite teams as part of job requirement
  • Proven good communication skills
  • 3+ years of experience in Elastic APM
  • 3 years with Scala
  • 3 years with Kubernetes Clusters

*All Telecommuters will be required to adhere to UnitedHealth Group’s Telecommuter Policy.

Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you’ll find a far-reaching choice of benefits and incentives. The salary for this role will range from $72,800 to $130,000 annually based on full-time employment. We comply with all minimum wage laws as applicable.

Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone–of every race, gender, sexuality, age, location, and income–deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups, and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes — an enterprise priority reflected in our mission.

UnitedHealth Group is an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.

UnitedHealth Group is a drug - free workplace. Candidates are required to pass a drug test before beginning employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer - Remote
Principal Site Reliability Engineer - Remote

UnitedHealth Group • Eden Prairie (MN)

Hybrid
Confidential
Remote work options
Senior Software Engineer
Senior Software Engineer

Optum • Basking Ridge (NJ)

On-site
USD 92,000 - 164,000
Comprehensive benefits
Equity stock purchase plan
401k contribution
+1
Lead Cloud Engineer - AI Ops - Remote
Lead Cloud Engineer - AI Ops - Remote

Optum • Minnetonka (MN)

Hybrid
USD 113,000 - 193,000
Comprehensive benefits package
Incentive and recognition programs
Equity stock purchase
+1
Senior Software Engineer
Senior Software Engineer

UnitedHealth Group • Minnetonka (MN), Northern (KY)

Hybrid
Confidential
PTO & Holidays
Medical & Dental Coverage
401(k) plan
+5
Senior Software Engineer
Senior Software Engineer

Optum • Minnetonka (MN)

On-site
USD 92,000 - 164,000
Paid Time Off
Medical Plan
Dental, Vision, Life & AD&D Insurance
+6
Senior Software Engineer - Remote Nationwide or Hybrid in MN or DC
Senior Software Engineer - Remote Nationwide or Hybrid in MN or DC

Optum • Eden Prairie (MN)

Hybrid
USD 92,000 - 164,000
Benefits package
Equity stock purchase
401k contribution
Senior Software and DevOps Engineer - Remote
Senior Software and DevOps Engineer - Remote

Optum • Minnetonka (MN)

Hybrid
USD 92,000 - 164,000
Benefits package
Equity stock purchase
401(k)
Lead Site Reliability Engineer, Chief Digital Office
Lead Site Reliability Engineer, Chief Digital Office

Worky • Eden Prairie (MN)

Hybrid
USD 113,000 - 193,000
Remote work
Senior Software Engineer
Senior Software Engineer

Worky • Frisco (TX)

On-site
USD 91,700 - 163,700
Comprehensive benefits package
Equity stock purchase
401k contribution
Director, Technology Delivery and AI Engineering - Remote
Director, Technology Delivery and AI Engineering - Remote

Optum • Eden Prairie (MN)

Hybrid
USD 135,000 - 231,000