Principal ML Engineer - Large-Scale Training Performance

Advanced Micro Devices, Inc.

San Jose (CA)

On-site

USD 130,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology company is seeking a Principal Machine Learning Engineer to join their Models and Applications team. This role focuses on improving training efficiency in distributed systems for large models using AMD GPUs. The ideal candidate will have a strong background in distributed training algorithms and ML frameworks like PyTorch or TensorFlow. Candidates should hold a Master’s or PhD in a relevant field and possess excellent programming and communication skills. The position is located in San Jose, CA or Bellevue, WA.

Qualifications

  • Experience with ML/DL frameworks such as PyTorch, JAX, or TensorFlow.
  • Experience with distributed training frameworks.
  • Excellent programming skills in Python or C++.

Responsibilities

  • Train large models to convergence on AMD GPUs.
  • Improve the end-to-end training pipeline performance.
  • Collaborate across teams with various groups.

Skills

Distributed training pipelines
Distributed training algorithms (Data Parallel, Tensor Parallel)
ML/DL frameworks (PyTorch, JAX, TensorFlow)
GPU kernel optimization
Python or C++ programming
Communication skills
Problem-solving skills

Education

Master's degree or PhD in Computer Science, AI, or Machine Learning

Tools

Megatron-LM
MaxText
TorchTitan

Job description

A leading technology company is seeking a Principal Machine Learning Engineer to join their Models and Applications team. This role focuses on improving training efficiency in distributed systems for large models using AMD GPUs. The ideal candidate will have a strong background in distributed training algorithms and ML frameworks like PyTorch or TensorFlow. Candidates should hold a Master’s or PhD in a relevant field and possess excellent programming and communication skills. The position is located in San Jose, CA or Bellevue, WA.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal ML Engineer: Large-Scale Training & Performance
Principal ML Engineer: Large-Scale Training & Performance

AMD • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Principal ML Engineer: Large-Scale Training Performance
Principal ML Engineer: Large-Scale Training Performance

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 120,000 - 180,000
Comprehensive benefits package
Innovative work culture
Principal ML Engineer - Large Scale Training Performance Optimization
Principal ML Engineer - Large Scale Training Performance Optimization

AMD • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Comprehensive benefits package
Flexible work arrangements
Collaborative work environment
Senior ML Performance Engineer - Distributed Training
Senior ML Performance Engineer - Distributed Training

Odyssey • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Principal ML Engineer - Large Scale Training Performance Optimization
Principal ML Engineer - Large Scale Training Performance Optimization

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 130,000 - 160,000
Principal ML Engineer - Large Scale Training Performance Optimization
Principal ML Engineer - Large Scale Training Performance Optimization

Advanced Micro Devices • San Jose (CA)

On-site
USD 120,000 - 180,000
Comprehensive benefits package
Innovative work culture
Senior ML Engineer: Distributed Training on Neuron
Senior ML Engineer: Distributed Training on Neuron

Amazon • Seattle (WA)

On-site
USD 151,300 - 261,500
Comprehensive medical benefits
Flexible work-life balance
Mentorship opportunities
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements