ML Inference Systems Engineer — High-Performance Accelerator

Arago Inc.

Paris

Sur place

EUR 110 000 - 170 000

Plein temps

Il y a 5 jours
Soyez parmi les premiers à postuler

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Stock options
Health insurance
Pension contributions
Professional development

Résumé du poste

Arago Inc. in Paris is seeking a senior ML inference engineer to optimize execution on a custom accelerator. You will work across kernels, model execution, and multi-device distribution to maximize performance and efficiency.

You will contribute to the software stack, with emphasis on CUDA/Triton-based kernels, operator fusion, and advanced inference-serving techniques; strong C++/Python skills and English proficiency are required.

Qualifications

  • Strong experience in high-performance ML inference.
  • Deep understanding of computer architecture and memory hierarchies.
  • Experience developing and optimizing custom kernels (CUDA/Triton/ROCm).
  • Experience with operator fusion, tiling, scheduling and data movement optimization.
  • Experience with distributed model execution and parallelism.
  • Hands-on with modern inference-serving systems (e.g., KV-cache, continuous batching).
  • Strong C++ and Python skills; exposure to MLIR is a plus.
  • English proficiency.

Responsabilités

  • Analyze modern AI workloads to identify bottlenecks on Arago's accelerator.
  • Develop and optimize custom kernels and execution strategies to maximize device utilization.
  • Map models and operators across multiple Arago devices with efficient communication.
  • Develop inference-serving techniques like continuous batching and KV caches.
  • Build profiling and performance-analysis tools for kernels, models, and serving workloads.
  • Collaborate with hardware, compiler, and runtime teams to influence future features.

Connaissances

ML inference
GPU/accelerator
Kernel optimization
Distributed execution
Inference-serving
C++
Python
English

Description du poste

Arago Inc. in Paris is seeking a senior ML inference engineer to optimize execution on a custom accelerator. You will work across kernels, model execution, and multi-device distribution to maximize performance and efficiency.

You will contribute to the software stack, with emphasis on CUDA/Triton-based kernels, operator fusion, and advanced inference-serving techniques; strong C++/Python skills and English proficiency are required.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

ML Inference Systems Engineer - Accelerator Performance
ML Inference Systems Engineer - Accelerator Performance

Arago • Paris

Sur place
EUR 120 000 - 180 000
Stock options
Healthcare coverage
Pension contributions
+2
ML Inference Engineer | Accelerate AI on Custom Accelerator
ML Inference Engineer | Accelerate AI on Custom Accelerator

Arago • Paris

Sur place
EUR 90 000 - 130 000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 120 000 - 180 000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 90 000 - 130 000
Stock options
Healthcare coverage
Pension contributions
+2
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • France

Sur place
EUR 90 000 - 150 000
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 70 000
AI Compiler Engineer: Optimize ML on Custom Accelerators
AI Compiler Engineer: Optimize ML on Custom Accelerators

IC Resources • Grenoble

Sur place
EUR 90 000 - 120 000
Senior AI/ML Scientist: Quantization & AI Hardware
Senior AI/ML Scientist: Quantization & AI Hardware

Arago • Paris

Sur place
EUR 80 000 - 110 000
Competitive cash compensation
Stock options
Relocation bonus
+4
Senior Computing Architect for AI Accelerators (Remote EU)
Senior Computing Architect for AI Accelerators (Remote EU)

Axelera AI • Paris

Hybride
EUR 70 000 - 90 000
Attractive compensation package
Pension plan
Employee insurances
+1