ML Inference Engineer | Accelerate AI on Custom Accelerator

Arago

Paris

On-site

EUR 90,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Stock options
Healthcare coverage
Pension contributions
Professional development
25 days PTO

Job summary

Arago is seeking an experienced engineer to optimize AI model execution on its custom accelerator. You will work across kernels, model execution, multi-device distribution, and inference serving, shaping the software stack around the hardware capabilities.

You will contribute to high-performance ML inference, kernel optimization, and distributed execution while collaborating with hardware, compiler, and runtime teams to drive performance improvements.

Qualifications

  • Strong experience in high-performance ML inference and accelerator programming.
  • Deep understanding of computer architecture, memory hierarchies, and parallelism.
  • Experience developing and optimizing custom kernels using low-level environments (CUDA, Triton, ROCm/HIP).
  • Experience with operator fusion, tiling, scheduling, data movement optimization, and graph execution.

Responsibilities

  • Analyze modern AI workloads to identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago's accelerator.
  • Develop and optimize custom kernels and fused operators for maximum device utilization.
  • Design efficient mappings of models across multiple Arago devices and manage communication/synchronization.
  • Develop inference-serving techniques such as continuous batching and KV caches.
  • Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels and full models.

Skills

ML inference
GPU programming
Computer architecture
CUDA/Triton/ROCm
Kernels optimization
Distributed execution
Inference serving
C++
Python
MLIR
English proficiency

Tools

CUDA
TensorRT
ROCm
MLIR

Job description

Arago is seeking an experienced engineer to optimize AI model execution on its custom accelerator. You will work across kernels, model execution, multi-device distribution, and inference serving, shaping the software stack around the hardware capabilities.

You will contribute to high-performance ML inference, kernel optimization, and distributed execution while collaborating with hardware, compiler, and runtime teams to drive performance improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference Systems Engineer — High-Performance Accelerator
ML Inference Systems Engineer — High-Performance Accelerator

Arago Inc. • Paris

On-site
EUR 110,000 - 170,000
Stock options
Health insurance
Pension contributions
+1
ML Inference Systems Engineer - Accelerator Performance
ML Inference Systems Engineer - Accelerator Performance

Arago • Paris

On-site
EUR 120,000 - 180,000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago Inc. • Paris

On-site
EUR 110,000 - 170,000
Stock options
Health insurance
Pension contributions
+1
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

On-site
EUR 120,000 - 180,000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

On-site
EUR 90,000 - 130,000
Stock options
Healthcare coverage
Pension contributions
+2
AI Accelerator Architect & Performance Modeler
AI Accelerator Architect & Performance Modeler

Arago Inc. • Paris

On-site
EUR 150,000 - 210,000
Stock option plan
Healthcare coverage
Pension contributions
+1
Lead Software Architect for High-Performance AI Accelerators
Lead Software Architect for High-Performance AI Accelerators

Arago Inc. • Paris

On-site
EUR 110,000 - 150,000
Stock options
Healthcare
Pension contributions
+2
AI Accelerator Architecture Modeling Engineer
AI Accelerator Architecture Modeling Engineer

arago • Paris

On-site
EUR 90,000 - 120,000
Stock options
Healthcare coverage
Pension contributions
+3
Performance Software Engineer - AI Accelerator Kernels
Performance Software Engineer - AI Accelerator Kernels

Arago • Paris

On-site
EUR 110,000 - 170,000
Stock options
Relocation support
Healthcare
+3
AI Accelerator Software Team Lead
AI Accelerator Software Team Lead

Arago • Paris

On-site
EUR 110,000 - 170,000
Stock options
Healthcare + pension
25 days PTO