Training / AI Infrastructure Engineering & Research Paris

Genesis

Paris

Sur place

EUR 120 000 - 180 000

Plein temps

14 jours+
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Résumé du poste

Genesis in Paris, France is seeking a senior ML infrastructure engineer to accelerate large-scale foundation model training by profiling and eliminating bottlenecks across the stack, from data pipelines to GPU kernels.

You will design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, focusing on scalability, robustness, and high utilization, with deep expertise in CUDA/cuDNN/Triton and production-grade Python.

Qualifications

  • 8+ years in distributed systems, ML infra, or high-performance computing.
  • Production-grade Python development experience.
  • Expert in CUDA/cuDNN/Triton and kernel optimization.
  • Experience with PyTorch training jobs using data, context, pipeline, and model parallelism.
  • System-level mindset for hardware-software tuning and scaling.

Responsabilités

  • Profile bottlenecks and reduce wall-clock time across the training stack.
  • Design, build, and optimize distributed PyTorch training on multi-node GPU clusters.
  • Implement low-level CUDA/CuDNN/Triton kernels and integrate into training frameworks.
  • Develop monitoring and debugging tools for large-scale runs and regressions.
  • Tune CPU/GPU balance, memory, data throughput, and networking for maximum utilization.

Connaissances

Python
PyTorch
CUDA/CuDNN
Distributed systems

Outils

CUDA
cuDNN
Triton

Description du poste

What You'll Do
  • Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels

  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization

  • Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks

  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking

  • Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures

What You'll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years)

  • Production-grade expertise in Python

  • Low-level performance mastery: CUDA/cuDNN/Triton, CPU-GPU interactions, data movement, and kernel optimization

  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism

  • System-level mindset with a track record of tuning hardware-software interactions for maximum utilization

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior AI Training Systems Engineer
Senior AI Training Systems Engineer

Genesis • Paris

Sur place
EUR 120 000 - 180 000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • France

Sur place
EUR 110 000 - 160 000
Compute Infrastructure Lead Onsite (Paris, France)
Compute Infrastructure Lead Onsite (Paris, France)

S27a • Paris

Hybride
EUR 120 000 - 160 000
Compute Infrastructure Lead
Compute Infrastructure Lead

UMA • Paris

Sur place
EUR 120 000 - 190 000
Research Engineer, Forge
Research Engineer, Forge

Jobtailor • Paris

Sur place
EUR 70 000 - 100 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA • Courbevoie

Sur place
EUR 120 000 - 170 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA Gruppe • Courbevoie

Sur place
EUR 120 000 - 160 000
Senior Engineer, AI Performance, Embedded Systems
Senior Engineer, AI Performance, Embedded Systems

Jobtailor • Massy

Sur place
EUR 90 000 - 130 000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA Corporation • Courbevoie

Sur place
EUR 120 000 - 180 000
Developer Technology Engineer, Energy
Developer Technology Engineer, Energy

NVIDIA • France

Sur place
EUR 90 000 - 130 000