Senior Platform Architect: HPC & Kubernetes at Scale

EPAM Systems

New York (NY)

Hybrid

USD 145,000 - 170,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical,Dental and Vision Insurance
Health Savings Account
401(k) Retirement Savings Plan
Paid Time Off
Holiday Pay

Job summary

EPAM Systems, Inc. in New York seeks an experienced Infrastructure Engineer to operate and scale Kubernetes platforms across multi-cloud providers, supporting thousands of GPUs and large-scale workloads.

You will design and maintain HPC infrastructure, build production-grade tooling, and define SLIs/SLOs with incident response responsibilities, collaborating across Networking, Storage, Security, and AI/ML teams.

Qualifications

  • 10+ years of infrastructure engineering, cloud platforms, or HPC environments.
  • Expert-level Kubernetes experience at meaningful scale with cluster lifecycle, scheduling, and upgrades.
  • Advanced Python skills; Go/Rust/C++ are a strong plus.
  • Daily proficiency with Terraform for infrastructure as code.
  • Working knowledge of core AWS services (EC2, S3, EFS, FSx for Lustre).
  • Strong SRE background with monitoring, alerting, and incident response practices.
  • CKA/CKS certificates are highly preferred.

Responsibilities

  • Operate and scale Kubernetes platforms (EKS, GKE, and other distributions) including cluster lifecycle management, node pool optimization, and networking policies during periods of rapid growth.
  • Provision and manage HPC infrastructure through CI/CD pipelines spanning AWS, CoreWeave, GCP, OCI, and additional cloud providers.
  • Design and maintain job scheduling systems that efficiently allocate GPU compute resources across training and inference workloads.
  • Define SLIs/SLOs, build robust monitoring and alerting systems, and actively participate in incident response and post-incident reviews.
  • Develop production-quality tooling and automation to support multi-cloud infrastructure operations at scale.
  • Collaborate daily with Networking, Storage, Security, and AI/ML platform teams to ensure seamless cross-functional infrastructure delivery.

Skills

Kubernetes at scale
Python
Terraform
AWS knowledge
Site reliability engineering

Tools

Go
Rust
C++
CI/CD pipelines

Job description

EPAM Systems, Inc. in New York seeks an experienced Infrastructure Engineer to operate and scale Kubernetes platforms across multi-cloud providers, supporting thousands of GPUs and large-scale workloads.

You will design and maintain HPC infrastructure, build production-grade tooling, and define SLIs/SLOs with incident response responsibilities, collaborating across Networking, Storage, Security, and AI/ML teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Kubernetes Architect - Multi-Cloud Platform
Senior HPC Kubernetes Architect - Multi-Cloud Platform

EPAM Systems • United States

Remote
USD 150,000 - 230,000
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)

EPAM Systems • United States

Remote
USD 140,000 - 230,000
Lead Platform Engineer/Architect - HPC, Kubernetes
Lead Platform Engineer/Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 150,000 - 230,000
Platform Architect - HPC, Kubernetes
Platform Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 140,000 - 230,000
Senior Cloud DevOps Platform Lead (AWS & Kubernetes)
Senior Cloud DevOps Platform Lead (AWS & Kubernetes)

EPAM Systems • United States

Remote
USD 120,000 - 180,000
Healthcare benefits
Paid time off and sick leave
LinkedIn Learning access
+2
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Senior Cloud Developer – Kubernetes & AI Platform
Senior Cloud Developer – Kubernetes & AI Platform

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 93,000 - 214,000
HPC (High-Performance Computing) Consultant @ Remote
HPC (High-Performance Computing) Consultant @ Remote

BURGEON IT SERVICES LLC • United States

Remote
USD 120,000 - 180,000
Lead Platform Engineer/Architect - HPC, Kubernetes
Lead Platform Engineer/Architect - HPC, Kubernetes

EPAM Systems • New York (NY)

Hybrid
USD 145,000 - 170,000
Medical,Dental and Vision Insurance
Health Savings Account
401(k) Retirement Savings Plan
+2
Senior Kubernetes Engineer – GPU & HPC Platform
Senior Kubernetes Engineer – GPU & HPC Platform

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 130,000 - 180,000
Company-Paid lunch stipend
Medical, dental, vision benefits
401(k) matching up to 6%
+1