Senior/Lead AI/ML Performance Engineer

EPAM Systems Inc

United States

Remote

USD 150,000 - 230,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

EPAM Systems is seeking a Senior/Lead AI/ML Performance Engineer to benchmark and profile AI/ML workloads on TPU hardware, generating empirical data for the optimization solver's cost model. You will conduct JAX/XLA benchmarking on TPU slices and analyze TPU profiler traces to build hardware coefficient matrices.

The role involves profiling LLM training and inference performance, collaborating with optimization engineers to calibrate solver inputs and improve accuracy and efficiency of cost

Qualifications

  • Strong understanding of the modern AI/ML landscape and LLM architectures.
  • Hands-on experience with LLM models/pipelines (training and inference).
  • Proficiency in Python and/or Golang.
  • Familiarity with TensorFlow, PyTorch, and LangChain.
  • Experience with JAX/XLA and TPU-specific profiling tools.

Responsibilities

  • Conduct JAX/XLA benchmarking on physical TPU slices.
  • Capture and analyze TPU Profiler traces.
  • Build hardware coefficient matrices for use in the optimizer's cost model.
  • Profile LLM training/inference performance (FLOPS, memory access, token throughput).
  • Collaborate with Optimization Engineers to calibrate solver inputs.

Skills

AI/ML landscape knowledge
LLM architectures
Python
Golang

Tools

TensorFlow
PyTorch
LangChain
JAX/XLA
TPU profiling tools

Job description

We are looking for a Senior/Lead AI/ML Performance Engineer to benchmark and profile AI/ML workloads on TPU hardware to generate the empirical performance data that feeds the optimization solver's cost model.

Responsibilities
  • Conduct JAX/XLA benchmarking on physical TPU slices
  • Capture and analyze TPU Profiler traces
  • Build hardware coefficient matrices for use in the optimizer's cost model
  • Profile LLM training/inference performance (FLOPS, memory access, token throughput)
  • Collaborate with Optimization Engineers to calibrate solver inputs
Requirements
  • Strong understanding of the modern AI/ML landscape and LLM architectures
  • Hands-on experience with LLM models/pipelines (training and inference)
  • Proficiency in Python and/or Golang
  • Familiarity with TensorFlow, PyTorch, and LangChain
  • Experience with JAX/XLA and TPU-specific profiling tools
Nice to have
  • Experience with GPU/TPU performance benchmarking at scale
  • Background in ML systems or MLPerf-style benchmarking
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI/ML Performance Engineer - TPU Benchmarking
Senior AI/ML Performance Engineer - TPU Benchmarking

EPAM Systems Inc • United States

Remote
USD 150,000 - 230,000
Software Engineer III, TPU Performance, Hardware and Software Codesign
Software Engineer III, TPU Performance, Hardware and Software Codesign

Google • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Bonus target
Equity
Benefits
ML Performance Engineering Manager (TPU & Optimization)
ML Performance Engineering Manager (TPU & Optimization)

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

On-site
USD 180,000 - 320,000
Senior Optimization / Algorithms Engineer
Senior Optimization / Algorithms Engineer

EPAM Systems Inc • United States

On-site
USD 130,000 - 190,000
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure

Apple Inc. • San Francisco (CA)

On-site
USD 210,000 - 300,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
TPU Performance Engineer – ML Hardware & Compiler Co-Design
TPU Performance Engineer – ML Hardware & Compiler Co-Design

Google • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Bonus target
Equity
Benefits