Get more replies from employers
Send a job-specific resume in minutes.
cerence is seeking an experienced ML infrastructure engineer to design and operate distributed training systems for large neural networks across GPU clusters. You will optimize multi-node, multi-GPU execution and diagnose bottlenecks across compute, memory, and networking to improve training efficiency.
Ideal candidates have hands-on experience with distributed systems, PyTorch distributed training, and deep knowledge of GPU communication.
cerence is seeking an experienced ML infrastructure engineer to design and operate distributed training systems for large neural networks across GPU clusters. You will optimize multi-node, multi-GPU execution and diagnose bottlenecks across compute, memory, and networking to improve training efficiency.
Ideal candidates have hands-on experience with distributed systems, PyTorch distributed training, and deep knowledge of GPU communication.