Compute Infrastructure Lead

UMA

Paris

Sur place

EUR 120 000 - 190 000

Plein temps

Il y a 2 jours
Soyez parmi les premiers à postuler
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Résumé du poste

UMA is seeking a Compute Infrastructure Lead to own and scale the compute backbone that provisions, schedules, and runs training and data-processing workloads across multi-provider GPUs. You will ensure reliability, efficiency, and researcher velocity as the platform grows from research to production.

This hands-on role combines architecture, ops, and leadership. You will build elastic scheduling, distributed training pipelines, and observability, working closely with researchers and partners

Qualifications

  • 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering.
  • Proven track record building and operating infrastructure for large-scale AI model training.
  • Deep hands-on experience with GPU clouds and cluster ops.

Responsabilités

  • Own compute platform end to end—from provisioning GPU capacity to keeping training, eval, and processing jobs running at high utilization.
  • Build a multi-provider management layer to place and fail over workloads across cloud providers and HPC.
  • Design and operate the cloud scheduler with quotas, priority, and preemption.
  • Stand up a distributed compute framework for training/evaluation on heterogeneous hardware.
  • Orchestrate data-processing workloads at scale across CPU and GPUs.
  • Deliver virtualized dev sessions (VMs on the cluster) for interactive work.
  • Build observability across providers using Prometheus, Grafana, and model metrics.
  • Forecast demand, manage costs, and negotiate provider terms while maintaining reliability.
  • Help set production-grade practices as we move toward partner POCs and scaling.

Connaissances

GPU clusters
Distributed training
Python
Systems engineering
Observability
Cost optimization
High-performance networking
Vendor coordination

Outils

Slurm
Kubernetes
Ray
NCCL
InfiniBand
Grafana
Prometheus

Description du poste

Your Mission

As Compute Infrastructure Lead, you will own and scale the compute backbone of UMA: the systems that provision, schedule, and run training, evaluation, and data-processing workloads — reliably, efficiently, and at scale — so our models can go from research to production without the cluster becoming the bottleneck.

This is a hands‑on, high‑impact role. We already train on a dedicated GPU cluster with a working training stack and a strong team behind it, so you won't be starting from zero — but you'll have the mandate to shape the architecture that takes us from a research cluster to a production‑scale, multi‑provider fleet, and to production‑grade reliability as we start deploying POCs with industry partners. You'll take ownership of the compute platform end to end — multi‑provider capacity, scheduling, distributed training / eval / processing, virtualized developer environments, observability, cost, and the tooling researchers and engineers actually use — and, if that's where you want to go, grow into leading the compute infrastructure team as it scales.

The technical problem is unusually rich for this stage. We are de‑risking a stack built on pre‑training and online RL, then industrializing it: heterogeneous hardware (training GPUs, cheaper eval and processing GPUs, CPU), a real‑time learning loop, orchestrated data processing, interactive VMs on the cluster, and a fleet that will grow fast. Much of what we need — a true multi‑provider compute fabric with elastic scheduling, dynamic checkpointing, unified observability, and jobs that resume themselves — does not exist off the shelf. Data infrastructure (datasets, storage, versioning) is owned by a sister role; this role is compute, including how processing jobs actually run on it.

Key responsibilities :

  • Own our compute platform end to end — from provisioning GPU capacity across cloud providers to keeping training, eval, and processing jobs running at high utilization, with reliability, cost, and researcher velocity as first‑class goals
  • Build a multi‑provider management layer so we can place, burst, and fail over workloads across GPU clouds, hyperscalers, and HPC without rewriting jobs
  • Design and operate the cloud scheduler — quotas, priority, preemption, topology‑aware placement, and dynamic checkpointing so jobs survive node failure, preemption, and provider switches
  • Stand up a distributed compute framework for training and evaluation on heterogeneous hardware (e.g. Ray / similar), including the real‑time / online‑learning path
  • Orchestrate data‑processing workloads at scale — CPU and cheaper GPUs, batch and streaming — so post‑processing, dataset jobs, and training share one reliable compute fabric instead of ad‑hoc scripts
  • Deliver virtualized GPU/CPU dev sessions (VMs on the cluster) so engineers iterate interactively on the same hardware and software they train on, without burning dedicated boxes
  • Build observability that works from any provider — system metrics (Prometheus, Grafana), job traces and logs, and model metrics (e.g. MLflow) — so a hung NCCL job, a silent GPU, or a broken training curve is diagnosable in minutes, not days
  • Own capacity, cost, and provider relationships as a technical lead: forecast demand, pick the right mix of hardware and contracts, and help negotiate pricing and terms. This is not a commercial role — but procurement is part of making the infra succeed
  • Help set production‑grade practices (testing, reliability, fast iteration) as we move from R&D to partner POCs, and grow into leading the compute infra team if that's the path you want
What You Bring To The Table
  • 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering, at a senior, lead, or staff level
  • Proven track record building and operating infrastructure for large‑scale AI model training — not inference‑only. Multi‑node GPU clusters, distributed training (PyTorch / NCCL or equivalent), and keeping long‑running jobs healthy at scale
  • Deep, hands‑on experience with GPU clouds and cluster operations: provisioning, Linux, high‑performance networking (InfiniBand / RoCE), storage for training, utilization, and GPU/node failure modes
  • Built or owned schedulers and distributed frameworks (Slurm, Kubernetes, Ray, SkyPilot, or similar) — including checkpointing, elasticity, and preemption — so GPUs stay busy and jobs come back from failure
  • Treat observability, reliability, and cost as core engineering concerns, not afterthoughts
  • Strong Python and systems engineering, with the taste to build tooling that researchers actually want to use
  • Experience working with GPU providers on capacity and commercial terms — you can read a contract, push on price and availability, and still be the person who debugs the cluster at 2am
  • Ability to reason about systems end‑to‑end — performance, scalability, reliability, cost — and make and defend the right trade-offs
  • Thrive in a hands‑on, fast‑paced startup, building from a real (but small) cluster toward a production fleet: autonomous, rigorous, execution‑driven, easy to work with, and broadly curious about AI and systems
  • Bonus : online / continuous RL, real‑time training loops, or other always‑on learning systems
  • Bonus : multi‑cloud / multi‑provider fabrics (SkyPilot, similar), Ray / Anyscale, HPC centers, VM‑based GPU workstations / interactive cluster sessions, or standing up clusters from tens to hundreds of nodes
  • Bonus : robotics, autonomous vehicles, or other embodied/physical‑AI training stacks — adjacent large‑scale training (LLMs, multimodal, AV) counts strongly; robotics itself is not required
  • Bonus : public projects, open‑source contributions, maintained tools, or technical writing
  • We value exceptional builders over perfect resumes. If you have a world‑class track record building training infrastructure at scale and the drive to build the compute backbone that lets a robotics company scale, robotics experience is a plus, not a requirement.
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Compute Infrastructure Lead Onsite (Paris, France)
Compute Infrastructure Lead Onsite (Paris, France)

S27a • Paris

Hybride
EUR 120 000 - 160 000
Data & ML Infrastructure Lead
Data & ML Infrastructure Lead

S27a • Paris

Sur place
EUR 120 000 - 160 000
Data & ML Infrastructure Lead
Data & ML Infrastructure Lead

UMA • Paris

Sur place
EUR 70 000 - 100 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA • Courbevoie

Sur place
EUR 120 000 - 170 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA Gruppe • Courbevoie

Sur place
EUR 120 000 - 160 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA Corporation • Courbevoie

Sur place
EUR 120 000 - 180 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA AI • Paris

Sur place
EUR 120 000 - 180 000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • France

Sur place
EUR 110 000 - 160 000
ML Infrastructure Engineer
ML Infrastructure Engineer

Whitecircle • Paris

Sur place
EUR 90 000 - 140 000
Training / AI Infrastructure Engineering & Research Paris
Training / AI Infrastructure Engineering & Research Paris

Genesis • Paris

Sur place
EUR 120 000 - 180 000