Remote Senior AI Infra Architect: HPC, TPU/GPU Co-Design

EPAM Systems Inc

Chicago (IL)

Remote

USD 180,000 - 220,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

EPAM Systems, Inc. is seeking a Principal Architect to lead the architectural strategy for AI performance, optimization, and hardware-software co-design at scale.

You will define the vision for AI training and serving infrastructure, spearheading a Center of Excellence and guiding the roadmap across TPU/GPU architectures, ML models, and compiler toolchains. The role requires 15+ years in software engineering with a track record of leadership in kernel-level and distributed systems, bridging

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, or equivalent practical experience (Master's or Ph.D. preferred).
  • 15+ years of software engineering experience, with 8+ years focused on distributed systems, AI infrastructure, or high-performance computing (HPC) architecture.
  • 7+ years of experience designing and developing complex software systems in C++ or Python.
  • 5+ years of experience leading the architecture, design, and delivery of large-scale software products, frameworks, or developer ecosystems from inception to production.
  • Proven track record of architecting performance-critical systems at the kernel level, bridging hardware accelerators and high-level software frameworks.

Responsibilities

  • Define and drive the multi-year technical roadmap for high-performance AI kernels, custom operations, and hardware-software co-design targeting TPU and GPU architectures.
  • Scale and mentor a world-class technical practice, establishing architectural governance, engineering standards, and best practices across the organization.
  • Act as the principal technical liaison partnering with ML researchers, core framework architects (JAX, PyTorch), and compiler engineering teams (XLA, MLIR) to eliminate systemic bottlenecks and shape future hardware/software requirements.
  • Architect foundational infrastructure—including enterprise-grade benchmarking suites, automated autotuning frameworks, regression analysis pipelines, and comprehensive documentation—empowering the global developer community.
  • Anticipate industry shifts by tracking advancements in hardware architectures, emerging model topologies, and compiler innovations to unlock step-changes in AI training and inference efficiency.

Skills

C++
Python
Distributed systems
AI infrastructure
Kernel-level engineering

Education

Bachelor's degree in Computer Science or Electrical Engineering

Tools

JAX
PyTorch
XLA
MLIR
CUDA

Job description

EPAM Systems, Inc. is seeking a Principal Architect to lead the architectural strategy for AI performance, optimization, and hardware-software co-design at scale.

You will define the vision for AI training and serving infrastructure, spearheading a Center of Excellence and guiding the roadmap across TPU/GPU architectures, ML models, and compiler toolchains. The role requires 15+ years in software engineering with a track record of leadership in kernel-level and distributed systems, bridging

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Infrastructure & Accelerator Architect
Lead AI Infrastructure & Accelerator Architect

EPAM Systems Inc • San Jose (CA)

Remote
USD 180,000 - 220,000
Medical, Dental and Vision Insurance (
Health Savings Account
Flexible Spending Accounts (Healthcare
+5
AI Infrastructure & Hardware-Accelerator Architect
AI Infrastructure & Hardware-Accelerator Architect

EPAM Systems, Inc. • United States

Remote
USD 180,000 - 220,000
Health Insurance
401(k) Retirement Plan
Paid Time Off
Lead AI Infrastructure & Accelerator Architect
Lead AI Infrastructure & Accelerator Architect

EPAM Systems Inc • Seattle (WA)

Remote
USD 180,000 - 220,000
Medical, Dental and Vision Insurance (
Health Savings Account
401(k) Matching
+4
Senior AI Network Architect for GPU Clusters & DPUs
Senior AI Network Architect for GPU Clusters & DPUs

AMSYS Innovative Solutions, LLC • Houston (TX)

On-site
USD 150,000 - 190,000
Remote AI Architect & Principal Software Engineer
Remote AI Architect & Principal Software Engineer

EPAM Systems Inc • Conshohocken

Remote
USD 140,000 - 165,000
Health insurance
401(k) matching plan
Paid time off
Remote AI Architect & Principal Engineer
Remote AI Architect & Principal Engineer

EPAM Systems Inc • Los Angeles (CA)

Remote
USD 140,000 - 165,000
Medical, Dental and Vision Insurance
Health Savings Account
Flexible Spending Accounts (Healthcare
+12
Senior CPU Performance Architect for AI/HPC Systems
Senior CPU Performance Architect for AI/HPC Systems

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior AI Architecture Lead
Senior AI Architecture Lead

EPAM Systems Inc • Conshohocken

Remote
USD 160,000 - 230,000
Medical Insurance
401(k) Plan
Paid Time Off
+3
Principal Multi-GPU AI-HPC System Architect
Principal Multi-GPU AI-HPC System Architect

NVIDIA • United States

Hybrid
USD 184,000 - 287,500
Equity
Benefits
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits