ML Inference Systems Engineer - Accelerator Performance

Arago

Paris

Sur place

EUR 120 000 - 180 000

Plein temps

Il y a 7 jours
Soyez parmi les premiers à postuler

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Stock options
Healthcare coverage
Pension contributions
Professional development
25 days PTO

Résumé du poste

Arago in Paris is building a cutting-edge AI accelerator and seeks a senior engineer to optimize model execution and inference serving. You will work across kernels, runtimes, and multi-device distribution, shaping the software stack around Arago's hardware.

Ideal candidates have deep experience in high-performance ML inference, CUDA/Triton/ROCm, C++ and Python, and a proven track record with modern inference-serving systems. Fluency in English required; French is a plus.

Qualifications

  • Extensive experience in high-performance ML inference and accelerator programming.
  • Deep understanding of computer architecture, memory hierarchies and performance bottlenecks.
  • Experience developing custom kernels with CUDA, Triton, ROCm/HIP or similar.
  • Familiar with operator fusion, tiling, scheduling and data movement optimization.
  • Experience with distributed model execution and parallelism.
  • Hands-on experience with modern inference-serving systems (e.g., vLLM, TensorRT-LLM) including KV-cache management.
  • Strong C++ and Python skills, with MLIR exposure a plus.
  • Proficient English.

Responsabilités

  • Analyze modern AI workloads and identify bottlenecks at kernel, runtime, memory, and system levels.
  • Develop and optimize kernels, fused operators, and execution strategies to maximize device utilization.
  • Design efficient mappings of models and operators across multiple Arago devices, including communication and synchronization strategies.
  • Develop inference-serving techniques such as continuous batching, paged KV caches, and prefill/decode strategies.
  • Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads.
  • Collaborate with hardware, compiler, and runtime teams to influence future software abstractions.

Connaissances

High-performance ML inference
GPU/accelerator programming
Distributed model execution
C++ and Python
MLIR familiarity

Outils

CUDA
Triton
ROCm/HIP
MLIR dialects

Description du poste

Arago in Paris is building a cutting-edge AI accelerator and seeks a senior engineer to optimize model execution and inference serving. You will work across kernels, runtimes, and multi-device distribution, shaping the software stack around Arago's hardware.

Ideal candidates have deep experience in high-performance ML inference, CUDA/Triton/ROCm, C++ and Python, and a proven track record with modern inference-serving systems. Fluency in English required; French is a plus.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

ML Inference Systems Engineer — High-Performance Accelerator
ML Inference Systems Engineer — High-Performance Accelerator

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1
ML Inference Engineer | Accelerate AI on Custom Accelerator
ML Inference Engineer | Accelerate AI on Custom Accelerator

Arago • Paris

Sur place
EUR 90 000 - 130 000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 120 000 - 180 000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 90 000 - 130 000
Stock options
Healthcare coverage
Pension contributions
+2
Senior AI/ML Scientist: Quantization & AI Hardware
Senior AI/ML Scientist: Quantization & AI Hardware

Arago • Paris

Sur place
EUR 80 000 - 110 000
Competitive cash compensation
Stock options
Relocation bonus
+4
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • France

Sur place
EUR 90 000 - 150 000
Senior Signal Processing Engineer — AI Hardware & Optics
Senior Signal Processing Engineer — AI Hardware & Optics

Arago • Paris

Sur place
EUR 65 000 - 95 000
Healthcare coverage
Stock option plan
Relocation bonus
+2
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 70 000
Senior AI/ML Scientist
Senior AI/ML Scientist

Arago • Paris

Sur place
EUR 80 000 - 110 000
Competitive cash compensation
Stock options
Relocation bonus
+4