ML Infrastructure Engineer

United States Digital Space LLC

Paris (TX)

Hybrid

USD 125,000 - 182,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive salary + equity
Hybrid work from Paris with relocation
Top-tier medical insurance in France
L&D budget and hardware/tools
Team off-sites twice a year

Job summary

White Circle is seeking an ML Infrastructure Engineer to join its AI safety team focused on policy enforcement and optimization for AI systems. The role requires hands-on experience with distributed RL/post-training systems, Python, PyTorch or JAX, and GPU debugging, with relocation to Paris for a hybrid setup.

You will build scalable pipelines, data control systems, and inference infrastructures, while developing agentic toolchains and multi-agent environments for robust model iteration and

Qualifications

  • Hands-on experience designing and running distributed RL/post-training systems at scale.
  • Strong Python (concurrency, async, multiprocessing, performance optimization) and PyTorch or JAX.
  • Debugging distributed GPU workloads across CUDA, drivers, containers, NCCL, networking, storage, and checkpointing.
  • Profiling across the stack (py-spy, PyTorch profiler, Nsight, perf, tracing).
  • Inference stacks: vLLM, SGLang, TensorRT-LLM, Dynamo, or custom serving.
  • Ability to connect system metrics to model behavior and learning dynamics.
  • Relocation to Paris (hybrid) required.

Responsibilities

  • Build scalable RL and post-training pipelines, including smoke tuning runs for quality testing and ablations.
  • Design data control systems for rollouts, replay, filtering, evaluation, and policy updates.
  • Tune training and inference end-to-end for throughput: networking, memory, scheduling, data loading, storage, checkpointing.
  • Build infrastructure for model iteration (experiment runs, artifacts, evals, dashboards, reproducibility, cost visibility) and inference infrastructure for post-training and eval loops.
  • Build agentic development environments: coding-agent harnesses, tool integrations, runtime sandboxes, multi-agent orchestration.

Skills

Python concurrency
PyTorch
JAX
Distributed systems
CUDA
Profiling tools
Kubernetes
C++

Tools

NCCL
Nsight
PyTorch profiler
TensorRT-LLM

Job description

We're looking for an ML Infrastructure Engineer to join White Circle, an AI Safety company building the policy enforcement and optimization layer for AI systems. Backed by $11M from senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, and DeepMind, White Circle processes 100M+ API calls monthly and runs its own LLMs in production.

You will
  • Build scalable RL and post-training pipelines, including smoke tuning runs for quality testing and ablations.
  • Design data control systems for rollouts, replay, filtering, evaluation, and policy updates.
  • Tune training and inference end-to-end for throughput: networking, memory, scheduling, data loading, storage, checkpointing, I/O.
  • Build infrastructure for model iteration (experiment runs, artifacts, evals, dashboards, reproducibility, cost visibility) and inference infrastructure for post-training and eval loops.
  • Build agentic development environments: coding-agent harnesses, tool integrations, runtime sandboxes, multi-agent orchestration.
Requirements
  • Hands-on experience designing and running distributed RL/post-training systems at scale (rollouts, replay buffers, reward signals, policy updates, eval loops).
  • Strong Python (concurrency, async, multiprocessing, performance optimization) and PyTorch or JAX.
  • Debugging distributed GPU workloads across CUDA, drivers, containers, NCCL, networking, storage, and checkpointing.
  • Profiling across the stack (py-spy, PyTorch profiler, Nsight, perf, tracing).
  • Inference stacks: vLLM, SGLang, TensorRT-LLM, Dynamo, or custom serving.
  • Ability to connect system metrics to model behavior and learning dynamics.
  • Relocation to Paris (hybrid) required.
Bonus
  • Public builder footprint: open-source contributions to RL, distributed ML, inference, eval, or agent infra; active technical presence on X.
  • Experience at high-bar AI infra/research teams (xAI, Qwen, ByteDance, Prime Intellect, or similar).
  • Ownership of custom training frameworks, trainers, schedulers, or data loaders.
  • GPU clusters on Kubernetes, Slurm, Ray; NCCL, RDMA, InfiniBand, RoCE, or EFA.
  • Rust, C++, CUDA, or Go; serious use of agentic coding tools (Claude Code, Codex, or similar).
We offer
  • Competitive salary + equity.
  • Hybrid work from Paris with relocation package.
  • Top-tier medical insurance in France and flexible time off.
  • L&D budget, all hardware and tools you need, plus covered AI agent and IDE subscriptions.
  • Team off-sites twice a year.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer
ML Infrastructure Engineer

Moonfire • Paris (TX)

Hybrid
USD 115,000 - 173,000
Equity
Flexible time off
Relocation package
+4
ML Research Engineer
ML Research Engineer

United States Digital Space LLC • Paris (TX)

Hybrid
USD 103,000 - 137,000
Hybrid work in Paris/London
Relocation support for Paris
Private health insurance
+2
ML Infra Engineer - Scale RL Pipelines (Paris)
ML Infra Engineer - Scale RL Pipelines (Paris)

United States Digital Space LLC • Paris (TX)

Hybrid
USD 125,000 - 182,000
Competitive salary + equity
Hybrid work from Paris with relocation
Top-tier medical insurance in France
+2
Member of Technical Staff - ML Infrastructure Engineer, Post-training
Member of Technical Staff - ML Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
AIML Engineer
AIML Engineer

Qubeaxis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Performance bonus (up to 20% of base)
Equity participation
Health, dental, and vision insurance
+3
QA Engineer
QA Engineer

United States Digital Space LLC • Paris (TX)

Hybrid
USD 59,000 - 89,000
Hybrid work Paris/London
Relocation package to Paris
Medical insurance in France
+3
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Remote option
Visa sponsorship
Relocation support
+2
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior Machine Learning Engineer (LLMs)
Senior Machine Learning Engineer (LLMs)

Albiware Inc. • Chicago (IL)

On-site
USD 140,000 - 210,000
Competitive salary
Generous PTO
Medical, dental, and vision coverage
+2
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Equity incentives
Visa sponsorship
Relocation support
+2