Lead GPU ML Training Performance Engineer

Advanced Micro Devices, Inc.

San Jose (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Advanced Micro Devices, Inc. is seeking a driven performance engineer to optimize posttraining workloads on AMD Instinct GPUs in San Jose, CA. The role involves collaborating on complex issues across various teams to enhance efficiency, stability, and throughput in training pipelines.

The ideal candidate will have a strong foundation in deep learning, proficiency in GPU performance engineering, and experience with tools like PyTorch. You will work on optimizing multiGPU training solutions and delivering reproducible pipelines for internal and external developers.

Qualifications

  • Proven experience in GPU performance engineering for deep learning.
  • Hands-on with SFT, LoRA and RL-based training at scale.
  • Comfortable reading/writing kernels when needed.

Responsibilities

  • Lead performance for finetuning and RL training solutions on AMD GPUs.
  • Improve throughput, memory efficiency, and stability across data, model and optimizer steps.
  • Optimize multiGPU/multinode training and communication patterns.

Skills

GPU performance engineering
Deep learning frameworks (ROCm/HIP, Triton)
Strong PyTorch experience
Proficient in Python and C++
Experience with distributed systems

Education

B.S./M.S./Ph.D. in Computer Science, Computer Engineering, or Electrical Engineering

Job description

Advanced Micro Devices, Inc. is seeking a driven performance engineer to optimize posttraining workloads on AMD Instinct GPUs in San Jose, CA. The role involves collaborating on complex issues across various teams to enhance efficiency, stability, and throughput in training pipelines.

The ideal candidate will have a strong foundation in deep learning, proficiency in GPU performance engineering, and experience with tools like PyTorch. You will work on optimizing multiGPU training solutions and delivering reproducible pipelines for internal and external developers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead GPU Performance Engineer for AI Training and Finetuning
Lead GPU Performance Engineer for AI Training and Finetuning

AMD • San Jose (CA)

Hybrid
USD 120,000 - 160,000
Senior GPU Performance Engineer for AI Training
Senior GPU Performance Engineer for AI Training

CareerArc • San Jose (CA)

Hybrid
USD 150,000 - 200,000
Competitive salary
Comprehensive benefits
Principal / Senior GPU SW Performance Engineer — Post‑Training
Principal / Senior GPU SW Performance Engineer — Post‑Training

AMD • San Jose (CA)

On-site
USD 120,000 - 160,000
Lead Datacenter GPU Performance Architect
Lead Datacenter GPU Performance Architect

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 150,000 - 190,000
AI Systems & GPU Performance Engineer Intern
AI Systems & GPU Performance Engineer Intern

Advanced Micro Devices • San Jose (CA), Santa Clara (CA)

Hybrid
USD 22,000 - 45,000
Senior GPU Software Performance Engineer – Post-Training
Senior GPU Software Performance Engineer – Post-Training

AMD • San Jose (CA)

On-site
USD 170,000 - 250,000
Senior GPU/AI Systems Engineer - Performance & ML
Senior GPU/AI Systems Engineer - Performance & ML

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits
Senior AI Performance Architect — GPU & Network
Senior AI Performance Architect — GPU & Network

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Comprehensive benefits package
AI/ML Compiler Engineer, High-Performance GPUs
AI/ML Compiler Engineer, High-Performance GPUs

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 260,000
Principal / Senior GPU SW Performance Engineer — Post‑Training
Principal / Senior GPU SW Performance Engineer — Post‑Training

AMD • San Jose (CA)

On-site
USD 170,000 - 250,000