Senior HPC Kubernetes Architect - Multi-Cloud Platform

EPAM Systems

United States

Remote

USD 150,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

EPAM Systems Inc. is building and operating Kubernetes platforms at massive scale across multiple clouds to power thousands of GPUs. You will design, implement, and operate scalable infrastructure, tooling, and job scheduling to maximize GPU utilization and reliability.

We value expert-level Kubernetes, Python and IaC skills, with experience in AWS, GCP, and Terraform. You will collaborate with Networking, Storage, Security, and AI/ML teams to deliver reliable, scalable cloud infrastructure.

Qualifications

  • 10+ years of experience in infrastructure engineering, cloud platforms, or HPC environments.
  • Expert-level Kubernetes experience at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets.
  • Advanced Python skills with production-grade tooling; experience with Go, Rust, or C++ is a strong plus.
  • Daily proficiency in Terraform for writing and reviewing infrastructure as code.
  • Working knowledge of core AWS services including EC2, S3, EFS, and FSx for Lustre.
  • Strong site reliability engineering background with monitoring, alerting, and incident response practices.
  • CKA, CKS certificates are highly preferred.

Responsibilities

  • Operate and scale Kubernetes platforms (EKS, GKE, and other distributions) including cluster lifecycle management, node pool optimization, and networking policies during periods of rapid growth
  • Provision and manage HPC infrastructure through CI/CD pipelines spanning AWS, CoreWeave, GCP, OCI, and additional cloud providers
  • Design and maintain job scheduling systems that efficiently allocate GPU compute resources across training and inference workloads
  • Define SLIs/SLOs, build robust monitoring and alerting systems, and actively participate in incident response and post-incident reviews
  • Develop production-quality tooling and automation to support multi-cloud infrastructure operations at scale
  • Collaborate daily with Networking, Storage, Security, and AI/ML platform teams to ensure seamless cross-functional infrastructure delivery

Skills

Kubernetes
Python
Go
Rust
C++
Terraform
AWS
GCP
CI/CD
Linux

Tools

EKS
GKE
CI/CD pipelines
Monitoring tooling

Job description

EPAM Systems Inc. is building and operating Kubernetes platforms at massive scale across multiple clouds to power thousands of GPUs. You will design, implement, and operate scalable infrastructure, tooling, and job scheduling to maximize GPU utilization and reliability.

We value expert-level Kubernetes, Python and IaC skills, with experience in AWS, GCP, and Terraform. You will collaborate with Networking, Storage, Security, and AI/ML teams to deliver reliable, scalable cloud infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)

EPAM Systems • United States

Remote
USD 140,000 - 230,000
Senior Platform Architect: HPC & Kubernetes at Scale
Senior Platform Architect: HPC & Kubernetes at Scale

EPAM Systems • New York (NY)

Hybrid
USD 145,000 - 170,000
Medical,Dental and Vision Insurance
Health Savings Account
401(k) Retirement Savings Plan
+2
Lead Platform Engineer/Architect - HPC, Kubernetes
Lead Platform Engineer/Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 150,000 - 230,000
Platform Architect - HPC, Kubernetes
Platform Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 140,000 - 230,000
Senior Cloud DevOps Platform Lead (AWS & Kubernetes)
Senior Cloud DevOps Platform Lead (AWS & Kubernetes)

EPAM Systems • United States

Remote
USD 120,000 - 180,000
Healthcare benefits
Paid time off and sick leave
LinkedIn Learning access
+2
Senior Kubernetes Engineer – GPU & HPC Platform
Senior Kubernetes Engineer – GPU & HPC Platform

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 130,000 - 180,000
Company-Paid lunch stipend
Medical, dental, vision benefits
401(k) matching up to 6%
+1
HPC (High-Performance Computing) Consultant @ Remote
HPC (High-Performance Computing) Consultant @ Remote

BURGEON IT SERVICES LLC • United States

Remote
USD 120,000 - 180,000
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Senior Cloud Developer – Kubernetes & AI Platform
Senior Cloud Developer – Kubernetes & AI Platform

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 93,000 - 214,000
Senior Platform Engineer — GPU Cloud Orchestration
Senior Platform Engineer — GPU Cloud Orchestration

STN Incorporated • United States

Hybrid
USD 140,000 - 180,000