Staff Software Engineer, GPU Infra & ML Systems

Reflection AI Ltd

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided
Parental leave
Unlimited PTO (US)
Visa sponsorship

Job summary

Reflection AI Ltd. is seeking an experienced engineer to design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.

You will develop systems powering synthetic data generation and reinforcement learning pipelines at scale, and build platforms capable of serving thousands of GPUs. Ideal candidates have hands-on experience with production infrastructure, strong GPU optimization expertise, and familiarity with advanced inference

Qualifications

  • Experience deploying and operating large-scale GPU systems for inference or model serving.
  • Several years of hands-on experience building and running production infrastructure.
  • Strong understanding of GPU performance characteristics and optimization techniques.
  • Experience with modern inference frameworks such as SGLang or Megatron.
  • Familiarity with distributed reinforcement learning infrastructure or rollout generation systems.
  • Experience optimizing throughput for large-scale model execution workloads.
  • Experience with GPU kernels or low-level performance optimization.
  • Familiarity with infrastructure used for synthetic data pipelines or RL training workflows.
  • Experience debugging performance issues across GPU, networking, and distributed execution layers.

Responsibilities

  • Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.
  • Develop systems that power synthetic data generation and reinforcement learning pipelines at scale.
  • Build high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
  • Optimize throughput, latency, and GPU utilization for large language model inference and rollout workloads.
  • Build infrastructure that supports reinforcement learning pipelines, including large-scale rollout generation, evaluation, and policy improvement loops.
  • Work closely with research teams to support distributed RL workloads and large-scale model evaluation infrastructure.
  • Improve performance of model execution through kernel-level optimization, model parallelism strategies, and GPU runtime improvements.
  • Develop distributed systems that enable large-scale synthetic data generation and RL-driven training workflows.
  • Diagnose and resolve performance bottlenecks across inference runtimes, GPU kernels, networking, and distributed compute systems.

Skills

GPU deployment
Production infrastructure
Distributed RL
Performance optimization
Kernel-level optimization

Tools

Megatron
SGLang
LLM runtimes

Job description

Reflection AI Ltd. is seeking an experienced engineer to design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.

You will develop systems powering synthetic data generation and reinforcement learning pipelines at scale, and build platforms capable of serving thousands of GPUs. Ideal candidates have hands-on experience with production infrastructure, strong GPU optimization expertise, and familiarity with advanced inference

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Systems Engineer: Large-Scale Inference & RL
Senior GPU Systems Engineer: Large-Scale Inference & RL

Reflection • New York (NY)

On-site
USD 150,000 - 200,000
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
Daily meals provided
+2
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and dental coverage
Access to a 401(k) retirement plan
+3
Remote AI Systems Engineer — Scalable GPU Infra
Remote AI Systems Engineer — Scalable GPU Infra

Bright Vision Technologies • United States

Remote
USD 90,000 - 100,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave for all new parents
+2
Member of Technical Staff - Mid-Training Infra
Member of Technical Staff - Mid-Training Infra

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+4
Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000
Remote ML Infra Architect: Scalable GPU Training
Remote ML Infra Architect: Scalable GPU Training

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Staff Research Software Engineer: Open ML Infra & RL Training
Staff Research Software Engineer: Open ML Infra & RL Training

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 260,000
Top-tier compensation & equity
Stock options
Health & wellness benefits
+5
Senior GPU Cluster Architect for AI Infra at Scale
Senior GPU Cluster Architect for AI Infra at Scale

Partner Company • United States

Remote
USD 184,000 - 318,000
Medical insurance
Dental insurance
Vision insurance
+1