ML Systems Engineer: Low-Latency RL Training Infra
Periodic Labs
Menlo Park (CA)
On-site
USD 300,000 - 400,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Periodic Labs, located in Menlo Park, is looking for a Systems Engineer to bridge infrastructure and research in AI. In this hybrid role, you'll manage the systems layer for fast and efficient model training, engaging closely with researchers. Ideal candidates have a Bachelor's degree and experience in large-scale inference infrastructure and low-level systems programming. The role offers a compensation range of $300,000-$400,000 and visa sponsorship. Join us in pushing the boundaries of scientific discovery.
Qualifications
Experience with low-level systems programming involving RDMA and network stack optimization.
Familiarity with GPU cluster scheduling using tools like Ray, Slurm, or Kubernetes.
Proficient in writing and optimizing CUDA kernels for distributed training.
Responsibilities
Build scheduling for GB series GPUs to minimize latency and maximize utilization.
Implement direct S3 checkpoint streaming to address I/O bottlenecks.
Design zero-copy RDMA weight synchronization to maintain a low-latency RL feedback loop.
Engage with the SGLang, Megatron, and Ray communities to contribute and influence improvements.
Skills
Large-scale inference infrastructure
Low-level systems programming
GPU cluster scheduling and orchestration
CUDA kernels optimization
Profiling and benchmarking distributed ML systems
Checkpoint management and streaming
Open source ML infrastructure contributions
Algorithm-infrastructure co-design
Education
Bachelor’s degree or equivalent
Job description
Periodic Labs, located in Menlo Park, is looking for a Systems Engineer to bridge infrastructure and research in AI. In this hybrid role, you'll manage the systems layer for fast and efficient model training, engaging closely with researchers. Ideal candidates have a Bachelor's degree and experience in large-scale inference infrastructure and low-level systems programming. The role offers a compensation range of $300,000-$400,000 and visa sponsorship. Join us in pushing the boundaries of scientific discovery.