Senior AI Compute Infrastructure Lead - Multi-Provider GPUs

UMA

Paris

Sur place

EUR 120 000 - 190 000

Plein temps

Il y a 8 jours
Générateur de candidature

Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.

Passez les filtres ATS

Résumé du poste

UMA is seeking a Compute Infrastructure Lead to own and scale the compute backbone that provisions, schedules, and runs training and data-processing workloads across multi-provider GPUs. You will ensure reliability, efficiency, and researcher velocity as the platform grows from research to production.

This hands-on role combines architecture, ops, and leadership. You will build elastic scheduling, distributed training pipelines, and observability, working closely with researchers and partners

Qualifications

  • 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering.
  • Proven track record building and operating infrastructure for large-scale AI model training.
  • Deep hands-on experience with GPU clouds and cluster ops.

Responsabilités

  • Own compute platform end to end—from provisioning GPU capacity to keeping training, eval, and processing jobs running at high utilization.
  • Build a multi-provider management layer to place and fail over workloads across cloud providers and HPC.
  • Design and operate the cloud scheduler with quotas, priority, and preemption.
  • Stand up a distributed compute framework for training/evaluation on heterogeneous hardware.
  • Orchestrate data-processing workloads at scale across CPU and GPUs.
  • Deliver virtualized dev sessions (VMs on the cluster) for interactive work.
  • Build observability across providers using Prometheus, Grafana, and model metrics.
  • Forecast demand, manage costs, and negotiate provider terms while maintaining reliability.
  • Help set production-grade practices as we move toward partner POCs and scaling.

Connaissances

GPU clusters
Distributed training
Python
Systems engineering
Observability
Cost optimization
High-performance networking
Vendor coordination

Outils

Slurm
Kubernetes
Ray
NCCL
InfiniBand
Grafana
Prometheus

Description du poste

UMA is seeking a Compute Infrastructure Lead to own and scale the compute backbone that provisions, schedules, and runs training and data-processing workloads across multi-provider GPUs. You will ensure reliability, efficiency, and researcher velocity as the platform grows from research to production.

This hands-on role combines architecture, ops, and leadership. You will build elastic scheduling, distributed training pipelines, and observability, working closely with researchers and partners

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Compute Infra Lead: Scale Multi-Cloud GPU Clusters
Compute Infra Lead: Scale Multi-Cloud GPU Clusters

S27a • Paris

Hybride
EUR 120 000 - 160 000
Compute Infrastructure Lead
Compute Infrastructure Lead

UMA • Paris

Sur place
EUR 120 000 - 190 000
Senior GPU Platform Engineer for AI on Kubernetes
Senior GPU Platform Engineer for AI on Kubernetes

Ubisoft • Paris

Sur place
EUR 90 000 - 130 000
Senior Data Infra Engineer for Scalable AI Compute
Senior Data Infra Engineer for Scalable AI Compute

Mistral.ai • Paris

Hybride
EUR 90 000 - 140 000
Head of GPU Cloud & HPC Engineering
Head of GPU Cloud & HPC Engineering

Scaleway • Paris

Hybride
EUR 140 000 - 190 000
Hybrid work
Office spaces
Meal service
+3
Compute Infrastructure Lead Onsite (Paris, France)
Compute Infrastructure Lead Onsite (Paris, France)

S27a • Paris

Hybride
EUR 120 000 - 160 000
Senior AI Inference Architect: Multi-Node GPU Scale
Senior AI Inference Architect: Multi-Node GPU Scale

NVIDIA • Marseille

Sur place
EUR 120 000 - 180 000
Senior Platform Engineer: GPU ML Platform on GCP
Senior Platform Engineer: GPU ML Platform on GCP

Ubisoft Entertainment Inc. • Paris

Sur place
EUR 90 000 - 120 000
Internal e-learning platform
Game library access
Clubs and activities (choir, yoga, etc
+1
Senior GPU Platform Engineer for AI Services & Kubernetes
Senior GPU Platform Engineer for AI Services & Kubernetes

Ubisoft • Paris

Sur place
EUR 90 000 - 120 000
E-learning platform
Game library access
CSE discounts
+1
Senior GPU Platform Engineer (Kubernetes + MLOps)
Senior GPU Platform Engineer (Kubernetes + MLOps)

Ubisoft Paris Studio • Paris

Sur place
EUR 90 000 - 140 000
Internal learning platform
Game library access
CSE discounts