Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave for all new parents
Paid time off and relocation support
Daily meals and team celebrations

Job summary

A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid understanding of reinforcement learning technologies. Comprehensive healthcare benefits, parental leave, and daily meals are provided, along with competitive salary and equity packages.

Qualifications

  • Experience with high-throughput model inference and mid-training workloads.
  • Strong technical skills in optimizing GPU utilization for inference.
  • Familiarity with large-scale distributed systems and RL workflows.

Responsibilities

  • Design and build GPU infrastructure for model inference.
  • Develop synthetic data generation systems and RL pipelines.
  • Optimize latency and throughput for large language models.
  • Support distributed RL workloads and model evaluation.

Skills

Experience deploying and operating large-scale GPU systems
Hands-on experience building and running production infrastructure
Understanding GPU performance characteristics
Experience with modern inference frameworks
Familiarity with distributed reinforcement learning infrastructure
Experience optimizing throughput for model execution workloads
Experience with GPU kernels and performance optimization
Familiarity with infrastructure for synthetic data pipelines
Debugging performance issues across GPU and distributed layers

Job description

A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid understanding of reinforcement learning technologies. Comprehensive healthcare benefits, parental leave, and daily meals are provided, along with competitive salary and equity packages.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
ML Infra Engineer: Scale GPU Training & Inference
ML Infra Engineer: Scale GPU Training & Inference

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
Staff Engineer, GPU AI Inference & RL Infrastructure
Staff Engineer, GPU AI Inference & RL Infrastructure

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Senior GPU Inference Engine Engineer
Senior GPU Inference Engine Engineer

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Senior ML Infra Architect — Large-Scale GPU Training
Senior ML Infra Architect — Large-Scale GPU Training

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior GPU Systems Engineer: Large-Scale Inference & RL
Senior GPU Systems Engineer: Large-Scale Inference & RL

Reflection • New York (NY)

On-site
USD 150,000 - 200,000
Principal ML Infra Engineer - GPU Inference & C++ Systems
Principal ML Infra Engineer - GPU Inference & C++ Systems

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000