Get more replies from employers
Send a job-specific resume in minutes.
Reflection is seeking a role to build and operate distributed training systems powering frontier models in SF. You will work with research teams to design scalable training runs across thousands of GPUs, optimizing throughput and stability while advancing production-ready training pipelines.
This role emphasizes collaboration with ML researchers, debugging across GPU stacks, and improving communication and memory efficiency in large-scale training environments.
Reflection is seeking a role to build and operate distributed training systems powering frontier models in SF. You will work with research teams to design scalable training runs across thousands of GPUs, optimizing throughput and stability while advancing production-ready training pipelines.
This role emphasizes collaboration with ML researchers, debugging across GPU stacks, and improving communication and memory efficiency in large-scale training environments.