Platform Architect - HPC, Kubernetes

EPAM Systems Inc

United States

Remote

USD 140,000 - 230,000

Full time

38 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

EPAM Systems Inc. is seeking a Platform Architect to support a multi-cloud HPC platform spanning AWS, GCP, OCI and other providers. The role emphasizes Kubernetes-heavy operations, scalable infrastructure, and production-grade tooling for GPU workloads across thousands of GPUs.

You will work with cross-functional platform teams to define SLIs/SLOs, implement monitoring, and ensure reliable upgrades and incident response in a fast-growing environment.

Qualifications

  • 7+ years of experience in infrastructure engineering, cloud platforms or HPC.
  • Expertise in Kubernetes at meaningful scale with node pools, scheduler debugging, and rolling upgrades across large fleets.
  • Proficiency in Python for production-grade tools and Terraform for writing and reviewing infrastructure as code daily.

Responsibilities

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale, including cluster lifecycle and node pool management.
  • Provision HPC infrastructure through CI/CD across AWS, CoreWeave, GCP and OCI.
  • Manage job scheduling to allocate GPU compute across training and inference workloads.
  • Define and maintain SLIs/SLOs, build monitoring and alerting, and participate in incident reviews.
  • Develop tooling and automation in production-quality code.
  • Coordinate daily with Networking, Storage, Security and AI/ML platform teams

Skills

Kubernetes
Python
Terraform
AWS
CI/CD
SRE Operations

Tools

Go
Rust
C++
CI/CD tooling

Job description

We are seeking a Platform Architect to support a customer that develops and manages several HPC clusters across AWS, CoreWeave, GCP and other providers, operating several thousand GPUs today and scaling 10x. This role is Kubernetes-heavy and requires strong software engineering skills, operating multi-cloud platform infrastructure where misconfigurations or failed upgrades cost thousands of GPU-hours.

Responsibilities
  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale, including cluster lifecycle, node pool management, networking policy, and stability during rapid growth
  • Provision HPC infrastructure through CI/CD across AWS, CoreWeave, GCP and OCI, with more providers coming
  • Management of job scheduling to allocate GPU compute across training and inference workloads
  • Definition and maintenance of SLIs/SLOs, build monitoring and alerting, and participation in incident response and post-incident reviews
  • Development of tooling and automation in production-quality code
  • Coordination daily with the Networking, Storage, Security and AI/ML platform teams
Requirements
  • 7+ years of experience in infrastructure engineering, cloud platforms or HPC
  • Expertise in Kubernetes at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
  • Proficiency in Python for production-grade tools, not only scripts
  • Proficiency in Terraform for writing and reviewing infrastructure as code daily
  • Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre) English proficiency at B2 level or higher
Nice to have
  • Familiarity with Amazon Elastic Kubernetes Service, Google Kubernetes Engine, Google Cloud Platform
  • Background in High-performance computing (HPC), Slurm, Lustre, Amazon FSx
  • Skills in Go, Rust, C++
  • Familiarity with CI/CD
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)

EPAM Systems Inc • United States

Remote
USD 140,000 - 230,000
HPC Solution Architect
HPC Solution Architect

Coda Search│Staffing • Dallas (TX)

On-site
USD 120,000 - 160,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

Linuxcareers • San Francisco (CA), Northern (KY)

On-site
USD 210,000 - 270,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Senior Kubernetes Developer – GPU & AI Infrastructure
Senior Kubernetes Developer – GPU & AI Infrastructure

GTN Technical Staffing • Town of Texas (WI), Fort Worth (TX)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work arrangement
Remote work flexibility
Senior Platform Engineer
Senior Platform Engineer

STN Incorporated • United States

Hybrid
USD 140,000 - 180,000
GPU / HPC Consultant
GPU / HPC Consultant

Arke • United States

On-site
USD 120,000 - 180,000