Staff Engineer - Large-Scale GPU Inference & RL Infra

Visa Hunt

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness benefits
Meals provided in office
22 weeks parental leave
Unlimited vacation (US)
Visa sponsorship
Team events

Job summary

Reflection is a research lab focused on making intelligence open and accessible. We design, build, and operate GPU-heavy infrastructure for high-throughput model inference and mid-training workloads.

Join a team that powers synthetic data generation, RL pipelines, and distributed model evaluation across thousands of GPUs. Expect collaboration with research teams and a focus on performance tuning, kernel optimization, and scalable distributed systems.

Qualifications

  • Experience deploying and operating large-scale GPU systems for inference or model serving.
  • Several years of hands-on experience building and running production infrastructure.
  • Strong understanding of GPU performance characteristics and optimization techniques.
  • Experience with modern inference frameworks such as SGLang, Megatron, or similar high-performance LLM runtimes.
  • Familiarity with distributed reinforcement learning infrastructure or rollout generation systems.
  • Experience optimizing throughput for large-scale model execution workloads.
  • Experience working with GPU kernels or low-level performance optimization.

Responsibilities

  • Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.
  • Develop systems that power synthetic data generation and reinforcement learning pipelines at scale.
  • Build high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
  • Diagnose and resolve performance bottlenecks across inference runtimes, GPU kernels, networking, and distributed compute systems.
  • Work closely with research teams to support distributed RL workloads and large-scale model evaluation infrastructure.

Skills

GPU infrastructure
production infrastructure
GPU performance optimization
distributed RL infrastructure
inference frameworks
kernel-level optimization
distributed compute systems
throughput optimization
networking

Tools

Megatron
SGLang

Job description

Reflection is a research lab focused on making intelligence open and accessible. We design, build, and operate GPU-heavy infrastructure for high-throughput model inference and mid-training workloads.

Join a team that powers synthetic data generation, RL pipelines, and distributed model evaluation across thousands of GPUs. Expect collaboration with research teams and a focus on performance tuning, kernel optimization, and scalable distributed systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Software Engineer: Scalable RL Training Systems
Research Software Engineer: Scalable RL Training Systems

Jobzhr • New York (NY)

On-site
USD 180,000 - 240,000
Salary and equity
Stock options
Healthcare
+3
Senior GPU Systems Engineer: Large-Scale Inference & RL
Senior GPU Systems Engineer: Large-Scale Inference & RL

Reflection • New York (NY)

On-site
USD 150,000 - 200,000
Member of Technical Staff - Mid-Training Infra
Member of Technical Staff - Mid-Training Infra

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness benefits
+5
Staff Engineer, Compute Platform & GPU Infra
Staff Engineer, Compute Platform & GPU Infra

Visa Hunt • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff - Mid-Training Infra
Member of Technical Staff - Mid-Training Infra

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave for all new parents
+2
Senior Distributed ML Training Engineer
Senior Distributed ML Training Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 310,000
Stock options
Health/dental/vision insurance
Meals provided in office
+1
Member of Technical Staff - Mid-Training Infra
Member of Technical Staff - Mid-Training Infra

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Staff Engineer, GPU AI Inference & RL Infrastructure
Staff Engineer, GPU AI Inference & RL Infrastructure

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000