ML Infra Engineer: Scale RL Pipelines & Inference

Moonfire

Paris (TX)

On-site

USD 115,000 - 173,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Flexible time off
Relocation package
Medical insurance
Learning & development
Equipment & tools
Team off-sites

Job summary

White Circle is seeking an ML Infrastructure Engineer to build the systems behind LLM post-training, RL, evaluation, and inference workflows. You’ll work closely with researchers on GPUs, training loops, data control systems, and the infrastructure decisions affecting model learning and product quality.

You will design and tune end-to-end training and inference stacks, build agentic development environments, and ensure scalable, reproducible experiments.

Qualifications

  • Designed, built, or maintained distributed RL/post-training systems at scale.
  • Fluent in PyTorch or JAX.
  • Proficient in Python, including concurrency and performance optimization.
  • Experience debugging distributed GPU workloads (CUDA, containers, NCCL, networking).
  • Experience with profiling tools (py-spy, PyTorch profiler, Nsight, perf).
  • Familiarity with inference stacks (vLLM, SGLang, TensorRT-LLM, Dynamo).
  • Ability to reason how infrastructure affects learning and eval quality.

Responsibilities

  • Build robust, scalable RL and post-training pipelines for quality testing and ablations.
  • Design data control systems governing data visibility and training flows.
  • Tune training and inference end-to-end for high throughput (networking, memory, I/O).
  • Investigate how infra choices affect learning dynamics and evaluation quality.
  • Build infrastructure for model iteration: experiments, artifacts, evals, dashboards, reproducibility.
  • Develop inference infrastructure supporting post-training and evaluation loops.
  • Create agentic development environments: coding-agent harnesses, tool integrations, sandboxes.
  • Collaborate with team to plan steps, discuss tradeoffs, and maintain communication.

Skills

Python programming
Distributed systems
Performance optimization
Research collaboration

Tools

PyTorch
JAX
CUDA
Kubernetes
Slurm
Ray
NCCL
UCX

Job description

White Circle is seeking an ML Infrastructure Engineer to build the systems behind LLM post-training, RL, evaluation, and inference workflows. You’ll work closely with researchers on GPUs, training loops, data control systems, and the infrastructure decisions affecting model learning and product quality.

You will design and tune end-to-end training and inference stacks, build agentic development environments, and ensure scalable, reproducible experiments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale Pipelines, GPUs & Platforms (Equity)
ML Infra Engineer: Scale Pipelines, GPUs & Platforms (Equity)

Alexander Chapman • New York (NY)

On-site
USD 120,000 - 160,000
Equity
Health insurance
Dental & Vision
Tech Lead Manager- MLRE, ML Systems
Tech Lead Manager- MLRE, ML Systems

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Engineering Manager, ML Infrastructure & Systems
Engineering Manager, ML Infrastructure & Systems

Cursor • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: Scale GPU Training & Data Pipelines
ML Infra Engineer: Scale GPU Training & Data Pipelines

Humble Robotics • United States

On-site
USD 150,000 - 230,000
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
ML Data Infrastructure Engineer
ML Data Infrastructure Engineer

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infrastructure Engineer: Scalable Training and Deployment
ML Infrastructure Engineer: Scalable Training and Deployment

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4