Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Cerence Inc. is seeking an experienced ML Systems Engineer to design and run distributed training systems for large neural networks across GPU clusters.
You will optimize multi-node, multi-GPU execution to maximize throughput and tenacity, while diagnosing bottlenecks across compute, memory, and network. Responsibilities include productionising large‑model training pipelines with PyTorch Distributed, Megatron‑LM and DeepSpeed, and improving training stability at scale.
Cerence Inc. is seeking an experienced ML Systems Engineer to design and run distributed training systems for large neural networks across GPU clusters.
You will optimize multi-node, multi-GPU execution to maximize throughput and tenacity, while diagnosing bottlenecks across compute, memory, and network. Responsibilities include productionising large‑model training pipelines with PyTorch Distributed, Megatron‑LM and DeepSpeed, and improving training stability at scale.