Senior GPU ML Infra Engineer — Mid-Training & Inference
Reflection AI
San Francisco (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave for all new parents
Paid time off and relocation support
Daily meals and team celebrations
Job summary
A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid understanding of reinforcement learning technologies. Comprehensive healthcare benefits, parental leave, and daily meals are provided, along with competitive salary and equity packages.
Qualifications
Experience with high-throughput model inference and mid-training workloads.
Strong technical skills in optimizing GPU utilization for inference.
Familiarity with large-scale distributed systems and RL workflows.
Responsibilities
Design and build GPU infrastructure for model inference.
Develop synthetic data generation systems and RL pipelines.
Optimize latency and throughput for large language models.
Support distributed RL workloads and model evaluation.
Skills
Experience deploying and operating large-scale GPU systems
Hands-on experience building and running production infrastructure
Understanding GPU performance characteristics
Experience with modern inference frameworks
Familiarity with distributed reinforcement learning infrastructure
Experience optimizing throughput for model execution workloads
Experience with GPU kernels and performance optimization
Familiarity with infrastructure for synthetic data pipelines
Debugging performance issues across GPU and distributed layers
Job description
A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid understanding of reinforcement learning technologies. Comprehensive healthcare benefits, parental leave, and daily meals are provided, along with competitive salary and equity packages.