Principal AI/ML Platform Engineer - Remote

Optum

Eden Prairie (MN)

Remote

USD 165,000 - 282,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity stock purchase
401k contribution
Comprehensive benefits package
Remote work option

Job summary

Optum is seeking a Principal AI/ML Platform Engineer to lead the end-to-end reference architecture for multi-tenant AI compute platforms across hybrid OpenShift clusters and public-cloud AI services. You will drive network standards, governance, and automated lifecycle management to ensure scalable, secure AI workloads.

Responsibilities include cost modeling, model-serving standardization, workload placement decisions, and robust IAM controls.

Qualifications

  • Bachelor’s degree or 4+ years of equivalent software/platform engineering experience in lieu of a degree.
  • 10+ years of experience in infrastructure, DevOps, SRE, or ML platform engineering.
  • 5+ years of experience operating Kubernetes or OpenShift at scale in production bare-metal or enterprise cloud environments.
  • 3+ years of experience designing and managing accelerated-compute (AI/GPU) infrastructure utilizing NVIDIA GPU Operator, NFD, MIG/time-slicing, and DCGM.
  • 3+ years of experience architecting hybrid AI platforms spanning self-hosted IaaS and public-cloud PaaS managed AI services (e.g., Azure AI Foundry, AWS Bedrock, or GCP Vertex AI).
  • 3+ years of experience with HPC/AI networking technologies, including InfiniBand or RoCEv2, GPUDirect RDMA, and NCCL collective communication limits.
  • 3+ years of experience managing multi-cluster fleets using RHACM (or equivalent) and GitOps tooling (Argo CD or Flux).
  • 3+ years of experience implementing enterprise security and IAM controls for software workloads (RBAC, OIDC/OAuth, Vault secrets management, mTLS).

Responsibilities

  • Own the end-to-end reference architecture for multi-tenant AI compute platforms across hybrid on-premises bare-metal OpenShift AI clusters and public-cloud managed AI platforms (Azure AI Foundry, AWS Bedrock, GCP Vertex AI).
  • Set network, latency, and topology standards for distributed training (NVLink, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL), ensuring interconnect boundaries are strictly maintained.
  • Establish cluster governance, GitOps workflows (Argo CD), RHACM policies, and automated lifecycle management for bare-metal accelerated compute nodes.
  • Design cost and utilization models including capex amortization, accelerator-sharing strategies (MIG/time-slicing), and cost-per-token/training-run math to inform accelerator procurement roadmaps.
  • Standardize model-serving platforms (vLLM, KServe) and inference gateways, setting quantization policies, provenance review gates, and intelligent model routing rules.
  • Define workload placement frameworks to determine self-hosted versus managed cloud deployment based on data residency, latency, cost, and compliance requirements.
  • Design and enforce identity, access, and security controls for AI workloads and autonomous agents, including least-privilege RBAC, Vault secret management, short-lived credentials, and mTLS.
  • Establish enterprise Service Level Objectives (SLOs), disaster recovery plans, and upgrade strategies for OpenShift, OpenShift AI, GPU operators, drivers, and firmware.
  • Partner with cross-functional AI teams, LLM gateway engineers, privacy, and security stakeholders to ensure seamless integration and HIPAA compliance.

Skills

Kubernetes
OpenShift
RHACM
GitOps
Security controls

Education

Bachelor’s degree or equivalent experience

Tools

Kubernetes
OpenShift
ARgo CD

Job description

Improve the lives of others while Caring. Connecting. Growing together.

Job Description - Principal AI/ML Platform Engineer - Remote (2386255)

Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care’s most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.

Primary Responsibilities:
  • Own the end-to-end reference architecture for multi-tenant AI compute platforms across hybrid on-premises bare-metal OpenShift AI clusters and public-cloud managed AI platforms (Azure AI Foundry, AWS Bedrock, GCP Vertex AI)
  • Set network, latency, and topology standards for distributed training (NVLink, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL), ensuring interconnect boundaries are strictly maintained
  • Establish cluster governance, GitOps workflows (Argo CD), RHACM policies, and automated lifecycle management for bare-metal accelerated compute nodes
  • Design cost and utilization models including capex amortization, accelerator-sharing strategies (MIG/time-slicing), and cost-per-token/training-run math to inform accelerator procurement roadmaps
  • Standardize model-serving platforms (vLLM, KServe) and inference gateways, setting quantization policies, provenance review gates, and intelligent model routing rules
  • Define workload placement frameworks to determine self-hosted versus managed cloud deployment based on data residency, latency, cost, and compliance requirements
  • Design and enforce identity, access, and security controls for AI workloads and autonomous agents, including least-privilege RBAC, Vault secret management, short-lived credentials, and mTLS
  • Establish enterprise Service Level Objectives (SLOs), disaster recovery plans, and upgrade strategies for OpenShift, OpenShift AI, GPU operators, drivers, and firmware
  • Partner with cross-functional AI teams, LLM gateway engineers, privacy, and security stakeholders to ensure seamless integration and HIPAA compliance

You’ll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.

Required Qualifications:
  • Bachelor’s degree or 4+ years of equivalent software/platform engineering experience in lieu of a degree
  • 10+ years of experience in infrastructure, DevOps, SRE, or ML platform engineering
  • 5+ years of experience operating Kubernetes or OpenShift at scale in production bare-metal or enterprise cloud environments
  • 3+ years of experience designing and managing accelerated-compute (AI/GPU) infrastructure utilizing NVIDIA GPU Operator, NFD, MIG/time-slicing, and DCGM
  • 3+ years of experience architecting hybrid AI platforms spanning self-hosted IaaS and public-cloud PaaS managed AI services (e.g., Azure AI Foundry, AWS Bedrock, or GCP Vertex AI)
  • 3+ years of experience with HPC/AI networking technologies, including InfiniBand or RoCEv2, GPUDirect RDMA, and NCCL collective communication limits
  • 3+ years of experience managing multi-cluster fleets using RHACM (or equivalent) and GitOps tooling (Argo CD or Flux)
  • 3+ years of experience implementing enterprise security and IAM controls for software workloads (RBAC, OIDC/OAuth, Vault secrets management, mTLS)
Preferred Qualifications:
  • Experience with distributed training frameworks (PyTorch DDP/FSDP, DeepSpeed, Ray, JAX) and batch scheduling systems (Kueue, Volcano)
  • Hands-on experience with LLM inference serving technologies (vLLM, TensorRT-LLM, KServe) and platform tooling such as OpenShift AI (RHOAI) or Kubeflow pipelines
  • Experience with high-performance parallel storage systems (Ceph/ODF, Lustre, IBM Storage Scale, VAST, WEKA)
  • Active Red Hat certifications (e.g., Red Hat Certified Architect / RHCA) or open-source contributions to CNCF, OpenShift, or AI infrastructure projects
  • Experience operating AI/ML platforms within regulated healthcare environments under HIPAA and UHG data privacy controls

*All employees working remotely will be required to adhere to UnitedHealth Group’s Telecommuter Policy.

Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you’ll find a far-reaching choice of benefits and incentives. The salary for this role will range from $164,600 - $282,200 annually based on full-time employment. We comply with all minimum wage laws as applicable.

Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.

At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.

UnitedHealth Group is an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.

UnitedHealth Group is a drug - free workplace. Candidates are required to pass a drug test before beginning employment.

UnitedHealth Group is committed to working with and providing reasonable accommodations to individuals with physical and mental disabilities. If you need special assistance or accommodation for any part of the application process, please call 1-866-566-8715 to be connected to Recruitment Services. Recruitment Services hours of operation are 7 a.m. to 7 p.m. CT, Monday through Friday.

UnitedHealth Group is a registered service mark of UnitedHealth Group, Inc. The UnitedHealth Group name with the dimensional logo, as well as the dimensional logo alone, are both service marks for the UnitedHealth Group, Inc.

Diversity creates a healthier atmosphere: UnitedHealth Group is an Equal Employment Opportunity/Affirmative Action employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, national origin, protected veteran status, disability status, sexual orientation, gender identity or expression, marital status, genetic information, or any other characteristic protected by law.

UnitedHealth Group is a drug-free workplace. Candidates are required to pass a drug test before beginning employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI/ML Platform Engineer - Remote
Principal AI/ML Platform Engineer - Remote

UnitedHealth Group • Eden Prairie (MN)

Remote
Confidential
Principal AI/ML Engineer
Principal AI/ML Engineer

Optum • Eden Prairie (MN)

Hybrid
USD 165,000 - 282,000
Lead AI/ML Engineer - Remote
Lead AI/ML Engineer - Remote

Optum • Eden Prairie (MN)

Remote
USD 146,000 - 250,000
Lead AI/ML Applied Scientist - Remote
Lead AI/ML Applied Scientist - Remote

Optum • Bellevue (WA)

Remote
USD 155,000 - 194,000
Lead SWE, AI Dev/IT Operations - Remote or Hybrid in DC or MN
Lead SWE, AI Dev/IT Operations - Remote or Hybrid in DC or MN

Optum • Minnetonka (MN)

Hybrid
USD 113,000 - 193,000
Comprehensive benefits package
Equity stock purchase program
401k contribution
Senior AI/ML Engineer - Remote
Senior AI/ML Engineer - Remote

Optum • Eden Prairie (MN)

Remote
USD 120,000 - 215,000
Lead AI/ML Engineer - Eden Prairie, MN or Schaumburg, IL
Lead AI/ML Engineer - Eden Prairie, MN or Schaumburg, IL

Optum • Eden Prairie (MN)

Hybrid
USD 146,000 - 250,000
Benefits package
Equity stock purchase
401k contribution
+1
Senior Software Engineer
Senior Software Engineer

Optum • Tempe (AZ)

Hybrid
USD 92,000 - 164,000
Comprehensive benefits package
Equity stock purchase
401k contribution
Director AI/ML Engineering - Eden Prairie, MN
Director AI/ML Engineering - Eden Prairie, MN

Optum • Eden Prairie (MN)

On-site
USD 177,000 - 303,000
Comprehensive benefits package
Equity stock purchase
401k contribution
Director, Technology Delivery and AI Engineering - Remote
Director, Technology Delivery and AI Engineering - Remote

UnitedHealth Group • Eden Prairie (MN)

Hybrid
Confidential
Comprehensive benefits
Equity stock purchase program
401k contribution