Compute Infra Lead: Scale Multi-Cloud GPU Clusters

S27a

Paris

Hybride

EUR 120 000 - 160 000

Plein temps

Il y a 13 jours
Générateur de candidature

Une candidature sur mesure pour ce poste — un CV personnalisé et une lettre de motivation qui correspondent directement à l’offre.

Passez les filtres ATS

Résumé du poste

UMA seeks a Compute Infrastructure Lead to own and scale the compute backbone that provisions, schedules, and runs training, evaluation, and data-processing workloads. You will shape a production-scale, multi-provider fleet and ensure reliability and researcher velocity.

You will lead end-to-end compute platform design, manage capacity and cost across cloud providers, and drive observability, tooling, and collaboration with researchers to ship production-ready systems.

Qualifications

  • 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering at a senior level.
  • Proven track record building and operating infrastructure for large-scale AI model training.
  • Hands-on experience with GPU clouds and cluster operations: provisioning, Linux, high-performance networking, storage for training.
  • Built or owned schedulers and distributed frameworks (Slurm, Kubernetes, Ray, SkyPilot) with checkpointing and elasticity.
  • Observability, reliability, and cost are core engineering concerns.

Responsabilités

  • Own our compute platform end to end—from provisioning GPU capacity to running training, eval, and processing jobs.
  • Build a multi-provider management layer to place workloads across GPU clouds and HPC.
  • Design and operate the cloud scheduler with quotas, priority, preemption, and dynamic checkpointing.
  • Stand up a distributed compute framework for training and evaluation on heterogeneous hardware.
  • Orchestrate data-processing workloads at scale to share one reliable compute fabric.
  • Deliver virtualized GPU/CPU dev sessions for interactive work on the same hardware used for training.
  • Build observability that works across providers (metrics, traces, logs, model metrics).
  • Own capacity, cost, and provider relationships as a technical lead.

Connaissances

8+ years exp
Python
GPU clouds
Linux
Observability
HPC

Outils

Slurm
Kubernetes
Ray
SkyPilot
NCCL

Description du poste

UMA seeks a Compute Infrastructure Lead to own and scale the compute backbone that provisions, schedules, and runs training, evaluation, and data-processing workloads. You will shape a production-scale, multi-provider fleet and ensure reliability and researcher velocity.

You will lead end-to-end compute platform design, manage capacity and cost across cloud providers, and drive observability, tooling, and collaboration with researchers to ship production-ready systems.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior AI Compute Infrastructure Lead - Multi-Provider GPUs
Senior AI Compute Infrastructure Lead - Multi-Provider GPUs

UMA • Paris

Sur place
EUR 120 000 - 190 000
Head of GPU Cloud Engineering
Head of GPU Cloud Engineering

Scaleway • Paris

Hybride
EUR 180 000 - 240 000
Compute Infrastructure Lead
Compute Infrastructure Lead

UMA • Paris

Sur place
EUR 120 000 - 190 000
GPU Cloud Engineering Leader
GPU Cloud Engineering Leader

Business At Work • Paris

Sur place
EUR 120 000 - 180 000
Hybrid work up to 3 days remote per wk
Meal service at HQ and regional sites
Well-being programs and gym access
+2
Head of GPU Cloud & HPC Engineering
Head of GPU Cloud & HPC Engineering

Scaleway • Paris

Hybride
EUR 140 000 - 190 000
Hybrid work
Office spaces
Meal service
+3
GPU Cloud Operations Lead — Hybrid, Global Impact
GPU Cloud Operations Lead — Hybrid, Global Impact

Webhosting • Paris

Hybride
EUR 110 000 - 150 000
Hybrid work
Modern offices
Healthy meals
+3
Engineering Team Lead, Bare-Metal Infra & GPU Compute
Engineering Team Lead, Bare-Metal Infra & GPU Compute

Mistral.ai • Paris

Sur place
EUR 110 000 - 150 000
Healthcare coverage
Relocation support
Wellness programs
Compute Infrastructure Lead Onsite (Paris, France)
Compute Infrastructure Lead Onsite (Paris, France)

S27a • Paris

Hybride
EUR 120 000 - 160 000
AI/HPC Cloud Infra Architect & DevOps Lead
AI/HPC Cloud Infra Architect & DevOps Lead

NVIDIA • Courbevoie

Sur place
EUR 120 000 - 170 000
Senior GPU Cloud Infra Architect & DevOps Lead
Senior GPU Cloud Infra Architect & DevOps Lead

NVIDIA Gruppe • Courbevoie

Sur place
EUR 120 000 - 160 000