Comprehensive medical, dental, and vision insurance
Fully paid parental leave
Paid time off
Daily meals provided
Job summary
Reflection, based in New York, is seeking an experienced professional to build and scale distributed training systems for frontier model pre-training. You will work closely with research teams to design large-scale training runs and optimize training efficiency across thousands of GPUs. The ideal candidate has experience with modern distributed training frameworks, strong debugging skills, and familiarity with GPU communication libraries. The company offers top-tier compensation, comprehensive health benefits, and a supportive team environment.
Qualifications
Experience building or operating distributed training systems for large machine learning models.
Strong experience with modern distributed training frameworks like Megatron, DeepSpeed.
Familiarity with large-scale model parallelism strategies.
Responsibilities
Build and scale distributed training systems for foundation models.
Design and operate large-scale training runs with research teams.
Optimize training throughput, stability, and efficiency.
Skills
Distributed training systems
Distributed training frameworks (Megatron, DeepSpeed)
Optimization of GPU utilization
Familiarity with NCCL
Debugging skills in GPU compute
Job description
Reflection, based in New York, is seeking an experienced professional to build and scale distributed training systems for frontier model pre-training. You will work closely with research teams to design large-scale training runs and optimize training efficiency across thousands of GPUs. The ideal candidate has experience with modern distributed training frameworks, strong debugging skills, and familiarity with GPU communication libraries. The company offers top-tier compensation, comprehensive health benefits, and a supportive team environment.