ML Systems Engineer: Low-Latency RL Training Infra

Periodic Labs

Menlo Park (CA)

On-site

USD 300,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Periodic Labs, located in Menlo Park, is looking for a Systems Engineer to bridge infrastructure and research in AI. In this hybrid role, you'll manage the systems layer for fast and efficient model training, engaging closely with researchers. Ideal candidates have a Bachelor's degree and experience in large-scale inference infrastructure and low-level systems programming. The role offers a compensation range of $300,000-$400,000 and visa sponsorship. Join us in pushing the boundaries of scientific discovery.

Qualifications

  • Experience with low-level systems programming involving RDMA and network stack optimization.
  • Familiarity with GPU cluster scheduling using tools like Ray, Slurm, or Kubernetes.
  • Proficient in writing and optimizing CUDA kernels for distributed training.

Responsibilities

  • Build scheduling for GB series GPUs to minimize latency and maximize utilization.
  • Implement direct S3 checkpoint streaming to address I/O bottlenecks.
  • Design zero-copy RDMA weight synchronization to maintain a low-latency RL feedback loop.
  • Engage with the SGLang, Megatron, and Ray communities to contribute and influence improvements.

Skills

Large-scale inference infrastructure
Low-level systems programming
GPU cluster scheduling and orchestration
CUDA kernels optimization
Profiling and benchmarking distributed ML systems
Checkpoint management and streaming
Open source ML infrastructure contributions
Algorithm-infrastructure co-design

Education

Bachelor’s degree or equivalent

Job description

Periodic Labs, located in Menlo Park, is looking for a Systems Engineer to bridge infrastructure and research in AI. In this hybrid role, you'll manage the systems layer for fast and efficient model training, engaging closely with researchers. Ideal candidates have a Bachelor's degree and experience in large-scale inference infrastructure and low-level systems programming. The role offers a compensation range of $300,000-$400,000 and visa sponsorship. Join us in pushing the boundaries of scientific discovery.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer – End-to-End Inference & RL
ML Systems Engineer – End-to-End Inference & RL

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 280,000
Startup equity
Health insurance
Competitive benefits
Staff ML Systems Engineer, Core Inference
Staff ML Systems Engineer, Core Inference

Together AI • San Francisco (CA)

On-site
USD 200,000 - 280,000
Startup equity
Health insurance
Competitive benefits
ML Systems Engineer: Post-Training for Next‑Gen Agent AI
ML Systems Engineer: Post-Training for Next‑Gen Agent AI

Scale AI, Inc. • New York (NY)

On-site
USD 250,000 - 350,000
Comprehensive health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
ML Systems Engineer — RL Training & Finetuning
ML Systems Engineer — RL Training & Finetuning

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Senior ML Infra Engineer — Real-Time, Low-Latency
Senior ML Infra Engineer — Real-Time, Low-Latency

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
ML Systems Engineer: Cloud‑Scale Training Infra
ML Systems Engineer: Cloud‑Scale Training Infra

Basis Research Institute • New York (NY)

On-site
USD 150,000 - 200,000
Competitive salary
Collaborative work environment
In-person events
ML Systems Engineer for RL & Inference Infrastructure
ML Systems Engineer for RL & Inference Infrastructure

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 160,000 - 210,000
AMD benefits
ML Systems Engineer: Distributed LLM Training & Inference
ML Systems Engineer: Distributed LLM Training & Inference

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 200,000 - 251,000
Comprehensive health coverage
Equity-based compensation
Retirement benefits
+3
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
AI Training Systems Engineer: Distributed & RL
AI Training Systems Engineer: Distributed & RL

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
Daily lunches and dinners provided
+1