Principal ML Engineer: Large-Scale Training & Performance

AMD

San Jose (CA)

Hybrid

USD 130,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive benefits package
Flexible work arrangements
Collaborative work environment

Job summary

A leading tech company is seeking a Principal Machine Learning Engineer in San Jose, CA. This role involves training large models on AMD GPUs and optimizing the training pipeline for efficiency. The ideal candidate should have extensive experience with distributed training algorithms and ML/DL frameworks like PyTorch and TensorFlow. The position offers the opportunity to work in a collaborative and innovative environment focused on AI advancements.

Qualifications

  • Experience with distributed training algorithms (Data Parallel, Tensor Parallel, etc.)
  • Familiarity with training large models at scale.
  • Experience optimizing GPU kernel performance.

Responsibilities

  • Train large models to convergence on AMD GPUs.
  • Improve the training pipeline performance.
  • Optimize distributed training pipeline and algorithm.
  • Contribute changes to open source.
  • Influence direction of AMD AI platform.

Skills

Distributed training pipelines
ML/DL frameworks (PyTorch, JAX, TensorFlow)
Communication skills
Problem-solving skills
Python or C++ programming

Education

Master's or PhD in Computer Science, AI, or ML

Tools

Megatron-LM
MaxText
TorchTitan

Job description

A leading tech company is seeking a Principal Machine Learning Engineer in San Jose, CA. This role involves training large models on AMD GPUs and optimizing the training pipeline for efficiency. The ideal candidate should have extensive experience with distributed training algorithms and ML/DL frameworks like PyTorch and TensorFlow. The position offers the opportunity to work in a collaborative and innovative environment focused on AI advancements.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal ML Engineer: Large-Scale Training Performance
Principal ML Engineer: Large-Scale Training Performance

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 120,000 - 180,000
Comprehensive benefits package
Innovative work culture
Principal ML Engineer - Large-Scale Training Performance
Principal ML Engineer - Large-Scale Training Performance

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 130,000 - 160,000
Senior GPU Training Performance Engineer
Senior GPU Training Performance Engineer

Advanced Micro Devices • San Jose (CA)

On-site
USD 120,000 - 160,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior GPU ML Training Performance Engineer
Senior GPU ML Training Performance Engineer

Advanced Micro Devices • San Jose (CA)

On-site
USD 130,000 - 170,000
Comprehensive health benefits
Inclusive workplace culture
Opportunities for career advancement
Senior ML Performance Engineer - Distributed Training
Senior ML Performance Engineer - Distributed Training

Odyssey • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Principal ML Engineer - Large Scale Training Performance Optimization
Principal ML Engineer - Large Scale Training Performance Optimization

AMD • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Comprehensive benefits package
Flexible work arrangements
Collaborative work environment
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Principal ML Engineer - Large Scale Training Performance Optimization
Principal ML Engineer - Large Scale Training Performance Optimization

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 130,000 - 160,000
Principal ML Engineer - Large Scale Training Performance Optimization
Principal ML Engineer - Large Scale Training Performance Optimization

Advanced Micro Devices • San Jose (CA)

On-site
USD 120,000 - 180,000
Comprehensive benefits package
Innovative work culture