Principal ML Engineer: Large-Scale Training & Performance
AMD
San Jose (CA)
Hybrid
USD 130,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Comprehensive benefits package
Flexible work arrangements
Collaborative work environment
Job summary
A leading tech company is seeking a Principal Machine Learning Engineer in San Jose, CA. This role involves training large models on AMD GPUs and optimizing the training pipeline for efficiency. The ideal candidate should have extensive experience with distributed training algorithms and ML/DL frameworks like PyTorch and TensorFlow. The position offers the opportunity to work in a collaborative and innovative environment focused on AI advancements.
Qualifications
Experience with distributed training algorithms (Data Parallel, Tensor Parallel, etc.)
Familiarity with training large models at scale.
Experience optimizing GPU kernel performance.
Responsibilities
Train large models to convergence on AMD GPUs.
Improve the training pipeline performance.
Optimize distributed training pipeline and algorithm.
Contribute changes to open source.
Influence direction of AMD AI platform.
Skills
Distributed training pipelines
ML/DL frameworks (PyTorch, JAX, TensorFlow)
Communication skills
Problem-solving skills
Python or C++ programming
Education
Master's or PhD in Computer Science, AI, or ML
Tools
Megatron-LM
MaxText
TorchTitan
Job description
A leading tech company is seeking a Principal Machine Learning Engineer in San Jose, CA. This role involves training large models on AMD GPUs and optimizing the training pipeline for efficiency. The ideal candidate should have extensive experience with distributed training algorithms and ML/DL frameworks like PyTorch and TensorFlow. The position offers the opportunity to work in a collaborative and innovative environment focused on AI advancements.