Get more replies from employers
Send a job-specific resume in minutes.
Cerence Inc. is seeking an ML Infrastructure Engineer to design and operate distributed training systems for large neural networks across GPU clusters.
You will optimize multi-node, multi-GPU execution and diagnose bottlenecks in compute, memory, and networking to maximize throughput and training stability. You will partner with research and applied ML teams to productionize large-model training pipelines using PyTorch Distributed, Megatron-LM, and DeepSpeed, with emphasis on scalable,
Cerence Inc. is seeking an ML Infrastructure Engineer to design and operate distributed training systems for large neural networks across GPU clusters.
You will optimize multi-node, multi-GPU execution and diagnose bottlenecks in compute, memory, and networking to maximize throughput and training stability. You will partner with research and applied ML teams to productionize large-model training pipelines using PyTorch Distributed, Megatron-LM, and DeepSpeed, with emphasis on scalable,