ML Inference Engineer | Accelerate AI on Custom Accelerator

Arago

Paris

Vor Ort

EUR 90.000 - 130.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, fertig zum Versenden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Stock options
Healthcare coverage
Pension contributions
Professional development
25 days PTO

Zusammenfassung

Arago is seeking an experienced engineer to optimize AI model execution on its custom accelerator. You will work across kernels, model execution, multi-device distribution, and inference serving, shaping the software stack around the hardware capabilities.

You will contribute to high-performance ML inference, kernel optimization, and distributed execution while collaborating with hardware, compiler, and runtime teams to drive performance improvements.

Qualifikationen

  • Strong experience in high-performance ML inference and accelerator programming.
  • Deep understanding of computer architecture, memory hierarchies, and parallelism.
  • Experience developing and optimizing custom kernels using low-level environments (CUDA, Triton, ROCm/HIP).
  • Experience with operator fusion, tiling, scheduling, data movement optimization, and graph execution.

Aufgaben

  • Analyze modern AI workloads to identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago's accelerator.
  • Develop and optimize custom kernels and fused operators for maximum device utilization.
  • Design efficient mappings of models across multiple Arago devices and manage communication/synchronization.
  • Develop inference-serving techniques such as continuous batching and KV caches.
  • Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels and full models.

Kenntnisse

ML inference
GPU programming
Computer architecture
CUDA/Triton/ROCm
Kernels optimization
Distributed execution
Inference serving
C++
Python
MLIR
English proficiency

Tools

CUDA
TensorRT
ROCm
MLIR

Jobbeschreibung

Arago is seeking an experienced engineer to optimize AI model execution on its custom accelerator. You will work across kernels, model execution, multi-device distribution, and inference serving, shaping the software stack around the hardware capabilities.

You will contribute to high-performance ML inference, kernel optimization, and distributed execution while collaborating with hardware, compiler, and runtime teams to drive performance improvements.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

ML Inference Systems Engineer — High-Performance Accelerator
ML Inference Systems Engineer — High-Performance Accelerator

Arago Inc. • Paris

Vor Ort
EUR 110.000 - 170.000
Stock options
Health insurance
Pension contributions
+1
ML Inference Systems Engineer - Accelerator Performance
ML Inference Systems Engineer - Accelerator Performance

Arago • Paris

Vor Ort
EUR 120.000 - 180.000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago Inc. • Paris

Vor Ort
EUR 110.000 - 170.000
Stock options
Health insurance
Pension contributions
+1
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Vor Ort
EUR 120.000 - 180.000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Vor Ort
EUR 90.000 - 130.000
Stock options
Healthcare coverage
Pension contributions
+2
AI Accelerator Architect & Performance Modeler
AI Accelerator Architect & Performance Modeler

Arago Inc. • Paris

Vor Ort
EUR 150.000 - 210.000
Stock option plan
Healthcare coverage
Pension contributions
+1
Lead Software Architect for High-Performance AI Accelerators
Lead Software Architect for High-Performance AI Accelerators

Arago Inc. • Paris

Vor Ort
EUR 110.000 - 150.000
Stock options
Healthcare
Pension contributions
+2
AI Accelerator Architecture Modeling Engineer
AI Accelerator Architecture Modeling Engineer

arago • Paris

Vor Ort
EUR 90.000 - 120.000
Stock options
Healthcare coverage
Pension contributions
+3
Performance Software Engineer - AI Accelerator Kernels
Performance Software Engineer - AI Accelerator Kernels

Arago • Paris

Vor Ort
EUR 110.000 - 170.000
Stock options
Relocation support
Healthcare
+3
AI Accelerator Software Team Lead
AI Accelerator Software Team Lead

Arago • Paris

Vor Ort
EUR 110.000 - 170.000
Stock options
Healthcare + pension
25 days PTO