Senior GPU Systems Engineer: Large-Scale Inference & RL
Reflection
New York (NY)
On-site
USD 150,000 - 200,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
Daily meals provided
Relocation support
Paid time off
Job summary
Reflection seeks a skilled professional in New York to design and operate large-scale GPU infrastructure for model inference and reinforcement learning. The role demands several years of experience in deploying GPU systems, optimizing model performance, and working with frameworks like SGLang and Megatron. The position offers competitive compensation, comprehensive benefits, and a supportive work environment with daily meals and team activities.
Qualifications
Several years of hands-on experience building and running production infrastructure.
Strong understanding of GPU performance characteristics and optimization techniques.
Experience deploying large-scale GPU systems for inference or model serving.
Responsibilities
Design and operate large-scale GPU infrastructure for model inference.
Develop systems for synthetic data generation and reinforcement learning pipelines.
Optimize GPU utilization for large language model inference.
Skills
GPU system deployment
Model serving
Optimizing throughput
Distributed reinforcement learning
Performance optimization techniques
Tools
SGLang
Megatron
Job description
Reflection seeks a skilled professional in New York to design and operate large-scale GPU infrastructure for model inference and reinforcement learning. The role demands several years of experience in deploying GPU systems, optimizing model performance, and working with frameworks like SGLang and Megatron. The position offers competitive compensation, comprehensive benefits, and a supportive work environment with daily meals and team activities.