Lead Platform Engineer/Architect - HPC, Kubernetes

EPAM Systems

United States

On-site

USD 150,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

EPAM Systems Inc. is building and operating Kubernetes platforms at massive scale across multiple clouds to power thousands of GPUs. You will design, implement, and operate scalable infrastructure, tooling, and job scheduling to maximize GPU utilization and reliability.

We value expert-level Kubernetes, Python and IaC skills, with experience in AWS, GCP, and Terraform. You will collaborate with Networking, Storage, Security, and AI/ML teams to deliver reliable, scalable cloud infrastructure.

Qualifications

  • 10+ years of experience in infrastructure engineering, cloud platforms, or HPC environments.
  • Expert-level Kubernetes experience at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets.
  • Advanced Python skills with production-grade tooling; experience with Go, Rust, or C++ is a strong plus.
  • Daily proficiency in Terraform for writing and reviewing infrastructure as code.
  • Working knowledge of core AWS services including EC2, S3, EFS, and FSx for Lustre.
  • Strong site reliability engineering background with monitoring, alerting, and incident response practices.
  • CKA, CKS certificates are highly preferred.

Responsibilities

  • Operate and scale Kubernetes platforms (EKS, GKE, and other distributions) including cluster lifecycle management, node pool optimization, and networking policies during periods of rapid growth
  • Provision and manage HPC infrastructure through CI/CD pipelines spanning AWS, CoreWeave, GCP, OCI, and additional cloud providers
  • Design and maintain job scheduling systems that efficiently allocate GPU compute resources across training and inference workloads
  • Define SLIs/SLOs, build robust monitoring and alerting systems, and actively participate in incident response and post-incident reviews
  • Develop production-quality tooling and automation to support multi-cloud infrastructure operations at scale
  • Collaborate daily with Networking, Storage, Security, and AI/ML platform teams to ensure seamless cross-functional infrastructure delivery

Skills

Kubernetes
Python
Go
Rust
C++
Terraform
AWS
GCP
CI/CD
Linux

Tools

EKS
GKE
CI/CD pipelines
Monitoring tooling

Job description

Join a high-growth infrastructure team operating Kubernetes platforms across multiple cloud providers at massive scale. You'll build the systems that power thousands of GPUs, where your code and configurations directly protect thousands of GPU-hours from costly failures. EPAM is where tech talent thrives building groundbreaking solutions, advancing your skills through world-class learning platforms, and working alongside a global community of problem-solvers to make the future real.

Req# 1103823066
Responsibilities
  • Operate and scale Kubernetes platforms (EKS, GKE, and other distributions) including cluster lifecycle management, node pool optimization, and networking policies during periods of rapid growth
  • Provision and manage HPC infrastructure through CI/CD pipelines spanning AWS, CoreWeave, GCP, OCI, and additional cloud providers
  • Design and maintain job scheduling systems that efficiently allocate GPU compute resources across training and inference workloads
  • Define SLIs/SLOs, build robust monitoring and alerting systems, and actively participate in incident response and post-incident reviews
  • Develop production-quality tooling and automation to support multi-cloud infrastructure operations at scale
  • Collaborate daily with Networking, Storage, Security, and AI/ML platform teams to ensure seamless cross-functional infrastructure delivery
Requirements
  • 10+ years of experience in infrastructure engineering, cloud platforms, or high-performance computing environments
  • Expert-level Kubernetes experience at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
  • Advanced Python skills with a track record of building production-grade tools, not just scripts; experience with Go, Rust, or C++ is a strong plus
  • Daily proficiency in Terraform for writing and reviewing infrastructure as code
  • Working knowledge of core AWS services including EC2, S3, EFS, and FSx for Lustre
  • Strong site reliability engineering background with experience building monitoring, alerting, and incident response practices
  • CKA, CKS certificates are highly preferred
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Architect - HPC, Kubernetes
Platform Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 140,000 - 230,000
Senior Platform Architect: HPC & Kubernetes at Scale
Senior Platform Architect: HPC & Kubernetes at Scale

EPAM Systems • New York (NY)

Hybrid
USD 145,000 - 170,000
Medical,Dental and Vision Insurance
Health Savings Account
401(k) Retirement Savings Plan
+2
Senior HPC Kubernetes Architect - Multi-Cloud Platform
Senior HPC Kubernetes Architect - Multi-Cloud Platform

EPAM Systems • United States

Remote
USD 150,000 - 230,000
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)

EPAM Systems • United States

Remote
USD 140,000 - 230,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
HPC (High-Performance Computing) Consultant @ Remote
HPC (High-Performance Computing) Consultant @ Remote

BURGEON IT SERVICES LLC • United States

Remote
USD 120,000 - 180,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Lead Platform Engineer/Architect - HPC, Kubernetes
Lead Platform Engineer/Architect - HPC, Kubernetes

EPAM Systems • New York (NY)

Hybrid
USD 145,000 - 170,000
Medical,Dental and Vision Insurance
Health Savings Account
401(k) Retirement Savings Plan
+2
Lead HPC and Systems Engineer
Lead HPC and Systems Engineer

EPAM Systems • United States

On-site
USD 150,000 - 190,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000