Staff Engineer - Large-Scale GPU Inference & RL Infra

reflectionai

San Francisco, New York (CA, NY)

On-site

USD 180,000 - 320,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided
Parental leave
Unlimited vacation
Visa sponsorship
Team events

Job summary

Reflection AI is building open, scalable AI infrastructure to enable high-throughput model inference and mid-training workloads. You will design and operate systems that support synthetic data generation, distributed RL, and evaluation across thousands of GPUs.

You will optimize throughput, latency, and GPU utilization while collaborating with research teams to advance large-scale model evaluation and deployment.

Qualifications

  • Experience deploying and operating large-scale GPU systems for inference or model serving.
  • Several years of hands-on experience building and running production infrastructure.
  • Strong understanding of GPU performance characteristics and optimization techniques.
  • Experience working with modern inference frameworks such as SGLang, Megatron, or similar high-performance LLM runtimes.
  • Familiarity with distributed reinforcement learning infrastructure or rollout generation systems.
  • Experience optimizing throughput for large-scale model execution workloads.
  • Experience working with GPU kernels or low-level performance optimization.
  • Familiarity with infrastructure used for synthetic data pipelines or RL training workflows.
  • Experience debugging performance issues across GPU, networking, and distributed execution layers.

Responsibilities

  • Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.
  • Develop systems that power synthetic data generation and reinforcement learning pipelines at scale.
  • Build high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
  • Diagnose and resolve performance bottlenecks across inference runtimes, GPU kernels, networking, and distributed compute systems.

Skills

GPU infrastructure
Model serving
RL pipelines
Distributed systems
Kernel optimization

Tools

SGLang
Megatron

Job description

Reflection AI is building open, scalable AI infrastructure to enable high-throughput model inference and mid-training workloads. You will design and operate systems that support synthetic data generation, distributed RL, and evaluation across thousands of GPUs.

You will optimize throughput, latency, and GPU utilization while collaborating with research teams to advance large-scale model evaluation and deployment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Systems Engineer: Large-Scale Inference & RL
Senior GPU Systems Engineer: Large-Scale Inference & RL

Reflection • New York (NY)

On-site
USD 150,000 - 200,000
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
Daily meals provided
+2
Senior RL Post-Training Systems Engineer (Equity)
Senior RL Post-Training Systems Engineer (Equity)

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
RL Post-Training Systems Architect (Equity Eligible)
RL Post-Training Systems Architect (Equity Eligible)

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Staff Engineer, GPU AI Inference & RL Infrastructure
Staff Engineer, GPU AI Inference & RL Infrastructure

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Senior RL Post-Training Frameworks Architect
Senior RL Post-Training Frameworks Architect

NVIDIA • United States

On-site
USD 180,000 - 260,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Distributed RL Systems Engineer — Scale Training & Inference
Distributed RL Systems Engineer — Scale Training & Inference

Luma AI • United States

Remote
USD 180,000 - 240,000
Staff Engineer: Distributed ML Training Systems
Staff Engineer: Distributed ML Training Systems

reflectionai • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 280,000
Stock options
Health insurance
Meals provided
+4