Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA

France

Sur place

EUR 90 000 - 150 000

Plein temps

Il y a 2 jours
Soyez parmi les premiers à postuler

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

NVIDIA in France seeks highly skilled software engineers to design and implement AI inference systems that scale large models with extreme efficiency. You will architect high‑performance inference stacks, optimize GPU kernels and compilers, and drive industry benchmarks for accelerated computing.

You will collaborate across inference, compiler, scheduling, and performance teams to push the pareto frontier of ML systems, enabling multi‑GPU, multi‑node, and multi‑cloud workloads.

Qualifications

  • Bachelor’s degree (or equivalent experience) in CS/CE/SE with 7+ years of experience; alternatively, Master’s degree in CS/CE/SE with 5+ years of experience; or PhD degree with the thesis and top-tier publications in ML Systems, GPU architecture, or high-performance computing.

Responsabilités

  • Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features; profile and optimize the inference framework (vLLM) with methods like speculative decoding, data/tensor/expert/pipeline-parallelism, prefill-decode disaggregation.
  • Develop, optimize, and benchmark GPU kernels (hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization; build and extend high-level DSLs and compiler infrastructure to boost kernel developer productivity while approaching peak hardware utilization.
  • Define and build inference benchmarking methodologies and tools; contribute both new benchmark and NVIDIA’s submissions to the industry-leading MLPerf Inference benchmarking suite.
  • Architect the scheduling and orchestration of containerized large-scale inference deployments on GPU clusters across clouds.
  • Conduct and publish original research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into NVIDIA’s software products.

Connaissances

Python
C/C++
Go
Rust
Parallel programming
Distributed systems
Deep learning

Formation

Bachelor's degree in CS/CE/SE
Master's degree in CS/CE/SE
PhD in ML Systems / GPU architecture / HPC

Outils

Docker
Kubernetes
Slurm
Nsight Systems/Compute
CUDA
MLIR/LLVM
Triton

Description du poste

NVIDIA in France seeks highly skilled software engineers to design and implement AI inference systems that scale large models with extreme efficiency. You will architect high‑performance inference stacks, optimize GPU kernels and compilers, and drive industry benchmarks for accelerated computing.

You will collaborate across inference, compiler, scheduling, and performance teams to push the pareto frontier of ML systems, enabling multi‑GPU, multi‑node, and multi‑cloud workloads.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior AI Solutions Architect for Large-Scale GPU HPC
Senior AI Solutions Architect for Large-Scale GPU HPC

NVIDIA AI • Aillas

Sur place
EUR 120 000 - 180 000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA AI • Aillas

Sur place
EUR 120 000 - 180 000
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • France

Sur place
EUR 90 000 - 150 000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • France

Sur place
EUR 110 000 - 160 000
AI Cloud GPU Solutions Architect for Datacentre Infra
AI Cloud GPU Solutions Architect for Datacentre Infra

NVIDIA • Courbevoie

Sur place
EUR 90 000 - 140 000
HPC & AI Computational Scientist - GPU Performance Engineer
HPC & AI Computational Scientist - GPU Performance Engineer

AMD • France

Sur place
EUR 90 000 - 130 000
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 70 000
Principal AI Architect: GPU & HPC Optimization Leader
Principal AI Architect: GPU & HPC Optimization Leader

NVIDIA • Courbevoie

Sur place
EUR 80 000 - 120 000
GPU HPC & AI Scientist — Paris Center of Excellence
GPU HPC & AI Scientist — Paris Center of Excellence

AMD • Paris

Sur place
EUR 70 000 - 100 000
ML Inference Systems Engineer — High-Performance Accelerator
ML Inference Systems Engineer — High-Performance Accelerator

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1