Senior Site Reliability Engineer

Optum

Hyderabad

On-site

INR 900,000 - 1,400,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Optum is seeking an experienced DevOps/SRE professional to design, deploy, and manage Azure infrastructure with AKS and robust IaC. You will build scalable CI/CD pipelines for data teams and enable self-service automation for data scientists.

You will monitor platform health, security, and performance while mentoring junior engineers and driving continuous improvement across the organization.

Qualifications

  • Bachelor's degree in CS/engineering or equivalent experience.
  • 5+ years in DevOps, SRE, or related role.
  • Experience with CI/CD tools (GitHub Actions).
  • Solid experience with data and AI platforms (Databricks, Snowflake).
  • Experience with orchestrators (Airflow, Data Factory).
  • Infrastructure as Code (Terraform, CloudFormation).
  • Azure cloud with AKS and core services.
  • Proficiency in scripting languages and monitoring tools.
  • Proven Kubernetes and containerization expertise.
  • Strong problem solving and collaboration skills.

Responsibilities

  • Design, deploy, and manage Azure infrastructure and services.
  • Build scalable, secure IaC with Terraform/CloudFormation.
  • Design, deploy and manage Kubernetes clusters (AKS).
  • Develop CI/CD pipelines for data engineering/science teams.
  • Create self-service automation for data scientists.
  • Monitor performance, security, and reliability of the platform.
  • Mentor juniors and collaborate across teams.
  • Implement observability with monitoring/logging solutions.
  • Drive continuous improvement and AI-enabled solutions.

Skills

Python
Bash
Ruby
Strong communication
Analytical skills
Problem solving
Team collaboration
Independent work
Vendor management
SRE mindset

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

GitHub Actions
Databricks
Snowflake
Airflow
Data Factory
Terraform
CloudFormation
Azure
AKS
Docker
Podman
Grafana
Prometheus
Kubernetes

Job description

Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.

Primary Responsibilities
  • Azure Cloud Experience:
    • Design, deploy, and manage Azure infrastructure and services
    • Optimize cloud resource utilization and cost management in Azure.
  • Infrastructure as Code (IaC):
    • Utilize IaC tools (such as Terraform, Ansible, or similar) to provision and manage infrastructure
    • Ensure infrastructure is scalable, secure, and resilient
  • Kubernetes Expertise:
    • Design, deploy, and manage Kubernetes clusters to support containerized applications
    • Implement and manage Kubernetes-based solutions for orchestration, scaling, and security
  • AI/DevOps/MLOps:
    • Design, build, and maintain robust, automated CI/CD pipelines for the data engineering and data science teams
    • Develop and manage tools that support these teams, ensuring a seamless and efficient experience for them
    • Utilize and build AI solutions to drive efficiencies across the platform engineering and wider teams
  • Automation Self-Service:
    • Implement automation for various operational processes, reducing manual intervention
    • Create self-service capabilities that empower Data Scientists to deploy and manage their applications independently
  • Platform Security Performance:
    • Ensure the security of the platform by implementing best practices and monitoring for vulnerabilities
    • Continuously monitor and optimize the performance of the platform to ensure high availability and reliability
  • Collaboration Mentorship:
    • Collaborate with cross-functional teams to align on project requirements and deliverables
    • Mentor junior team members, promoting best practices in DevOps and automation
  • Observability:
    • Implement monitoring and logging solutions to ensure system health and performance
    • Troubleshoot and resolve issues related to system performance, security, and reliability
  • Continuous Improvement:
    • Stay current with industry trends and advancements in DevOps practices and technologies
    • Identify opportunities for process improvements and drive initiatives to implement them
  • Design, develop, and deploy AI-powered solutions using no-code, low-code, and advanced platforms, translating business needs into scalable applications that enhance products, workflows, and decision-making
  • Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications
  • Bachelor''s degree in Computer Science, Engineering, or a related field (or equivalent work experience)
  • 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or a similar role
  • Experience with CI/CD tools preferably GitHub Actions
  • Solid experience of Data and AI platforms, preferably Databricks and Snowflake
  • Experience using orchestrating tools (Airflow, Data Factory)
  • Experience with Infrastructure as Code (Terraform, CloudFormation)
  • Expert knowledge of cloud platforms, with a focus on Azure
  • Familiarity with core Azure services - Storage, Networking, security, App Services, AKS
  • Proficiency in scripting languages (e.g., Python, Bash, Ruby)
  • Proficiency in monitoring tools (Splunk, Grafana, Prometheus)
  • Proven expertise in Kubernetes and containerization technologies (Docker, Podman)
  • Proven excellent problem-solving and analytical skills
  • Proven solid communication and collaboration skills
  • Proven ability to work independently and in a team-oriented, collaborative environment
  • Proven ability to work with third party vendors and support teams in a large multi disciplined organization
  • Skills:
    • Kubernetes, Docker and Containerization, CICD
    • Azure Cloud
    • Infrastructure as Code
    • Grafana and Prometheus
    • AKS
Preferred Qualifications
  • Experience in Azure, AWS, GCP and private cloud technologies
  • Experience using Firewalls, Load Balancers and DDOS solutions
  • Experience in a leadership or mentorship role
  • Knowledge of security best practices in DevOps

At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

UnitedHealth Group • Hyderabad

On-site
Confidential
Senior I O Engineering Consultant - Python, K8s, Azure cloud, Terraform
Senior I O Engineering Consultant - Python, K8s, Azure cloud, Terraform

Optum India • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Devops Engineer - SRE, Cloud Engineering
Devops Engineer - SRE, Cloud Engineering

Optum • Hyderabad

On-site
INR 2,500,000 - 4,500,000
DevOps Engineer
DevOps Engineer

Optum India • Dadri

On-site
INR 1,400,000 - 2,000,000
Senior Software Engineer I - Azure DevOps
Senior Software Engineer I - Azure DevOps

Optum India • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optum • Chennai District

On-site
INR 1,400,000 - 2,500,000
Senior Software Engineering Lead
Senior Software Engineering Lead

UnitedHealth Group • India

On-site
Confidential
Lead Full Stack Engineer - Data Structures and Algorithms
Lead Full Stack Engineer - Data Structures and Algorithms

Optum • Bengaluru

Hybrid
INR 3,500,000 - 6,500,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

UnitedHealth Group • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior Software Engineering Lead- Typescript, React, Terraform
Senior Software Engineering Lead- Typescript, React, Terraform

Optum India • Bengaluru

On-site
INR 4,000,000 - 7,000,000