ML Inference Systems Engineer - Accelerator Performance

Arago

Paris

On-site

EUR 120,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Stock options
Healthcare coverage
Pension contributions
Professional development
25 days PTO

Job summary

Arago in Paris is building a cutting-edge AI accelerator and seeks a senior engineer to optimize model execution and inference serving. You will work across kernels, runtimes, and multi-device distribution, shaping the software stack around Arago's hardware.

Ideal candidates have deep experience in high-performance ML inference, CUDA/Triton/ROCm, C++ and Python, and a proven track record with modern inference-serving systems. Fluency in English required; French is a plus.

Qualifications

  • Extensive experience in high-performance ML inference and accelerator programming.
  • Deep understanding of computer architecture, memory hierarchies and performance bottlenecks.
  • Experience developing custom kernels with CUDA, Triton, ROCm/HIP or similar.
  • Familiar with operator fusion, tiling, scheduling and data movement optimization.
  • Experience with distributed model execution and parallelism.
  • Hands-on experience with modern inference-serving systems (e.g., vLLM, TensorRT-LLM) including KV-cache management.
  • Strong C++ and Python skills, with MLIR exposure a plus.
  • Proficient English.

Responsibilities

  • Analyze modern AI workloads and identify bottlenecks at kernel, runtime, memory, and system levels.
  • Develop and optimize kernels, fused operators, and execution strategies to maximize device utilization.
  • Design efficient mappings of models and operators across multiple Arago devices, including communication and synchronization strategies.
  • Develop inference-serving techniques such as continuous batching, paged KV caches, and prefill/decode strategies.
  • Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads.
  • Collaborate with hardware, compiler, and runtime teams to influence future software abstractions.

Skills

High-performance ML inference
GPU/accelerator programming
Distributed model execution
C++ and Python
MLIR familiarity

Tools

CUDA
Triton
ROCm/HIP
MLIR dialects

Job description

Arago in Paris is building a cutting-edge AI accelerator and seeks a senior engineer to optimize model execution and inference serving. You will work across kernels, runtimes, and multi-device distribution, shaping the software stack around Arago's hardware.

Ideal candidates have deep experience in high-performance ML inference, CUDA/Triton/ROCm, C++ and Python, and a proven track record with modern inference-serving systems. Fluency in English required; French is a plus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference Systems Engineer — High-Performance Accelerator
ML Inference Systems Engineer — High-Performance Accelerator

Arago Inc. • Paris

On-site
EUR 110,000 - 170,000
Stock options
Health insurance
Pension contributions
+1
ML Inference Engineer | Accelerate AI on Custom Accelerator
ML Inference Engineer | Accelerate AI on Custom Accelerator

Arago • Paris

On-site
EUR 90,000 - 130,000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

On-site
EUR 120,000 - 180,000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago Inc. • Paris

On-site
EUR 110,000 - 170,000
Stock options
Health insurance
Pension contributions
+1
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

On-site
EUR 90,000 - 130,000
Stock options
Healthcare coverage
Pension contributions
+2
Lead Software Architect for High-Performance AI Accelerators
Lead Software Architect for High-Performance AI Accelerators

Arago Inc. • Paris

On-site
EUR 110,000 - 150,000
Stock options
Healthcare
Pension contributions
+2
Performance Software Engineer - AI Accelerator Kernels
Performance Software Engineer - AI Accelerator Kernels

Arago • Paris

On-site
EUR 110,000 - 170,000
Stock options
Relocation support
Healthcare
+3
AI Accelerator Architecture Modeling Engineer
AI Accelerator Architecture Modeling Engineer

arago • Paris

On-site
EUR 90,000 - 120,000
Stock options
Healthcare coverage
Pension contributions
+3
Lead Software Engineer, AI Accelerator Stack
Lead Software Engineer, AI Accelerator Stack

Arago • Paris

On-site
EUR 120,000 - 180,000
Stock options
Health insurance
Pension contributions
+2
AI Accelerator Architect & Performance Modeler
AI Accelerator Architect & Performance Modeler

Arago Inc. • Paris

On-site
EUR 150,000 - 210,000
Stock option plan
Healthcare coverage
Pension contributions
+1