Compute Infrastructure Lead Onsite (Paris, France)

S27a

Paris

Hybride

EUR 120 000 - 160 000

Plein temps

Il y a 8 jours
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Résumé du poste

UMA seeks a Compute Infrastructure Lead to own and scale the compute backbone that provisions, schedules, and runs training, evaluation, and data-processing workloads. You will shape a production-scale, multi-provider fleet and ensure reliability and researcher velocity.

You will lead end-to-end compute platform design, manage capacity and cost across cloud providers, and drive observability, tooling, and collaboration with researchers to ship production-ready systems.

Qualifications

  • 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering at a senior level.
  • Proven track record building and operating infrastructure for large-scale AI model training.
  • Hands-on experience with GPU clouds and cluster operations: provisioning, Linux, high-performance networking, storage for training.
  • Built or owned schedulers and distributed frameworks (Slurm, Kubernetes, Ray, SkyPilot) with checkpointing and elasticity.
  • Observability, reliability, and cost are core engineering concerns.

Responsabilités

  • Own our compute platform end to end—from provisioning GPU capacity to running training, eval, and processing jobs.
  • Build a multi-provider management layer to place workloads across GPU clouds and HPC.
  • Design and operate the cloud scheduler with quotas, priority, preemption, and dynamic checkpointing.
  • Stand up a distributed compute framework for training and evaluation on heterogeneous hardware.
  • Orchestrate data-processing workloads at scale to share one reliable compute fabric.
  • Deliver virtualized GPU/CPU dev sessions for interactive work on the same hardware used for training.
  • Build observability that works across providers (metrics, traces, logs, model metrics).
  • Own capacity, cost, and provider relationships as a technical lead.

Connaissances

8+ years exp
Python
GPU clouds
Linux
Observability
HPC

Outils

Slurm
Kubernetes
Ray
SkyPilot
NCCL

Description du poste

Your Mission

As Compute Infrastructure Lead, you will own and scale the compute backbone of UMA: the systems that provision, schedule, and run training, evaluation, and data-processing workloads — reliably, efficiently, and at scale — so our models can go from research to production without the cluster becoming the bottleneck.

This is a hands-on, high-impact role. We already train on a dedicated GPU cluster with a working training stack and a strong team behind it, so you won't be starting from zero — but you'll have the mandate to shape the architecture that takes us from a research cluster to a production-scale, multi-provider fleet, and to production-grade reliability as we start deploying POCs with industry partners. You'll take ownership of the compute platform end to end — multi-provider capacity, scheduling, distributed training / eval / processing, virtualized developer environments, observability, cost, and the tooling researchers and engineers actually use — and, if that's where you want to go, grow into leading the compute infrastructure team as it scales.

The technical problem is unusually rich for this stage. We are de-risking a stack built on pre-training and online RL, then industrializing it: heterogeneous hardware (training GPUs, cheaper eval and processing GPUs, CPU), a real-time learning loop, orchestrated data processing, interactive VMs on the cluster, and a fleet that will grow fast. Much of what we need — a true multi-provider compute fabric with elastic scheduling, dynamic checkpointing, unified observability, and jobs that resume themselves — does not exist off the shelf. Data infrastructure (datasets, storage, versioning) is owned by a sister role; this role is compute, including how processing jobs actually run on it.

Key responsibilities :

  • Own our compute platform end to end — from provisioning GPU capacity across cloud providers to keeping training, eval, and processing jobs running at high utilization, with reliability, cost, and researcher velocity as first-class goals

  • Build a multi-provider management layer so we can place, burst, and fail over workloads across GPU clouds, hyperscalers, and HPC without rewriting jobs

  • Design and operate the cloud scheduler — quotas, priority, preemption, topology-aware placement, and dynamic checkpointing so jobs survive node failure, preemption, and provider switches

  • Stand up a distributed compute framework for training and evaluation on heterogeneous hardware (e.g. Ray / similar), including the real-time / online-learning path

  • Orchestrate data-processing workloads at scale — CPU and cheaper GPUs, batch and streaming — so post-processing, dataset jobs, and training share one reliable compute fabric instead of ad-hoc scripts

  • Deliver virtualized GPU/CPU dev sessions (VMs on the cluster) so engineers iterate interactively on the same hardware and software they train on, without burning dedicated boxes

  • Build observability that works from any provider — system metrics (Prometheus, Grafana), job traces and logs, and model metrics (e.g. MLflow) — so a hung NCCL job, a silent GPU, or a broken training curve is diagnosable in minutes, not days

  • Own capacity, cost, and provider relationships as a technical lead: forecast demand, pick the right mix of hardware and contracts, and help negotiate pricing and terms. This is not a commercial role — but procurement is part of making the infra succeed

  • Help set production-grade practices (testing, reliability, fast iteration) as we move from R&D to partner POCs, and grow into leading the compute infra team if that's the path you want

What You Bring to the Table
  • 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering, at a senior, lead, or staff level

  • Proven track record building and operating infrastructure for large-scale AI model training — not inference-only. Multi-node GPU clusters, distributed training (PyTorch / NCCL or equivalent), and keeping long-running jobs healthy at scale

  • Deep, hands-on experience with GPU clouds and cluster operations: provisioning, Linux, high-performance networking (InfiniBand / RoCE), storage for training, utilization, and GPU/node failure modes

  • Built or owned schedulers and distributed frameworks (Slurm, Kubernetes, Ray, SkyPilot, or similar) — including checkpointing, elasticity, and preemption — so GPUs stay busy and jobs come back from failure

  • Treat observability, reliability, and cost as core engineering concerns, not afterthoughts

  • Strong Python and systems engineering, with the taste to build tooling that researchers actually want to use

  • Experience working with GPU providers on capacity and commercial terms — you can read a contract, push on price and availability, and still be the person who debugs the cluster at 2am

  • Ability to reason about systems end-to-end — performance, scalability, reliability, cost — and make and defend the right trade-offs

  • Thrive in a hands-on, fast-paced startup, building from a real (but small) cluster toward a production fleet: autonomous, rigorous, execution-driven, easy to work with, and broadly curious about AI and systems

  • Bonus : online / continuous RL, real-time training loops, or other always-on learning systems

  • Bonus : multi-cloud / multi-provider fabrics (SkyPilot, similar), Ray / Anyscale, HPC centers, VM-based GPU workstations / interactive cluster sessions, or standing up clusters from tens to hundreds of nodes

  • Bonus : robotics, autonomous vehicles, or other embodied/physical-AI training stacks — adjacent large-scale training (LLMs, multimodal, AV) counts strongly; robotics itself is not required

  • Bonus : public projects, open-source contributions, maintained tools, or technical writing

  • We value exceptional builders over perfect resumes. If you have a world-class track record building training infrastructure at scale and the drive to build the compute backbone that lets a robotics company scale, we strongly encourage you to apply — even if you don't tick every box. Robotics experience is a plus, not a requirement.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Compute Infrastructure Lead
Compute Infrastructure Lead

UMA • Paris

Sur place
EUR 120 000 - 190 000
Data & ML Infrastructure Lead
Data & ML Infrastructure Lead

S27a • Paris

Sur place
EUR 120 000 - 160 000
Data & ML Infrastructure Lead
Data & ML Infrastructure Lead

UMA • Paris

Sur place
EUR 70 000 - 100 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA • Courbevoie

Sur place
EUR 120 000 - 170 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA Gruppe • Courbevoie

Sur place
EUR 120 000 - 160 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA Corporation • Courbevoie

Sur place
EUR 120 000 - 180 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA AI • Paris

Sur place
EUR 120 000 - 180 000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • France

Sur place
EUR 110 000 - 160 000
ML Infrastructure Engineer
ML Infrastructure Engineer

Whitecircle • Paris

Sur place
EUR 90 000 - 140 000
HPC Engineer
HPC Engineer

Arlequin AI • Paris

Hybride
EUR 90 000 - 130 000