ML Systems Engineer — Inference Acceleration

Arago

Paris

Sur place

EUR 120 000 - 180 000

Plein temps

Il y a 6 jours
Soyez parmi les premiers à postuler

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Stock options
Healthcare coverage
Pension contributions
Professional development
25 days PTO

Résumé du poste

Arago in Paris is building a cutting-edge AI accelerator and seeks a senior engineer to optimize model execution and inference serving. You will work across kernels, runtimes, and multi-device distribution, shaping the software stack around Arago's hardware.

Ideal candidates have deep experience in high-performance ML inference, CUDA/Triton/ROCm, C++ and Python, and a proven track record with modern inference-serving systems. Fluency in English required; French is a plus.

Qualifications

  • Extensive experience in high-performance ML inference and accelerator programming.
  • Deep understanding of computer architecture, memory hierarchies and performance bottlenecks.
  • Experience developing custom kernels with CUDA, Triton, ROCm/HIP or similar.
  • Familiar with operator fusion, tiling, scheduling and data movement optimization.
  • Experience with distributed model execution and parallelism.
  • Hands-on experience with modern inference-serving systems (e.g., vLLM, TensorRT-LLM) including KV-cache management.
  • Strong C++ and Python skills, with MLIR exposure a plus.
  • Proficient English.

Responsabilités

  • Analyze modern AI workloads and identify bottlenecks at kernel, runtime, memory, and system levels.
  • Develop and optimize kernels, fused operators, and execution strategies to maximize device utilization.
  • Design efficient mappings of models and operators across multiple Arago devices, including communication and synchronization strategies.
  • Develop inference-serving techniques such as continuous batching, paged KV caches, and prefill/decode strategies.
  • Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads.
  • Collaborate with hardware, compiler, and runtime teams to influence future software abstractions.

Connaissances

High-performance ML inference
GPU/accelerator programming
Distributed model execution
C++ and Python
MLIR familiarity

Outils

CUDA
Triton
ROCm/HIP
MLIR dialects

Description du poste

Meet Arago and the Aragonians

Arago's mission is to reengineer the foundations of computing from first principles.

The explosive growth of AI is pushing the industry to rethink how processors are built. Arago is meeting that challenge with a proprietary technology that fuses optical and CMOS technologies to deliver an order-of-magnitude increase in performance.

Arago is the fastest, and currently the only, company to have built such a processor. It's backed by leading deep-tech investors and some of the most respected figures in semiconductors and computing, including the CEO of Arm, the founder of macOS who worked directly with Steve Jobs at Apple, an Nvidia Fellow, the Head of Optics at Google, and many other industry leaders.

Our work is guided by three clear values: do great things, move with high velocity, and operate as one unit. We work in a demanding environment where constant learning, ownership, and execution are expected, and where exceptional people have the opportunity to do their life’s work.

What you'll do

Optimize the execution and serving of modern AI models on Arago's custom accelerator. Work across kernels, model execution, multi-device distribution, runtime, and inference serving, while helping shape the software stack around the capabilities of Arago's hardware.

Required Skills and Experience
  • Strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering.

  • Deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks.

  • Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments.

  • Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads.

  • Strong understanding of distributed model execution, including tensor, pipeline, sequence, and/or expert parallelism and communication/computation overlap.

  • Hands-on experience with modern inference-serving systems such as vLLM, SGLang, TensorRT-LLM, or equivalent, including KV-cache management, continuous batching, paged attention, and prefill/decode scheduling.

  • Strong C++ and Python skills, and comfort working on a custom accelerator stack where compiler, runtime, kernels, and abstractions are actively being developed. Exposure to or experience with MLIR and MLIR dialects is a strong plus.

  • Language: English at a proficient level.

Responsibilities
  • Analyze modern AI workloads and identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago's accelerator.

  • Develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization.

  • Design efficient mappings of models and operators across multiple Arago devices, including communication and synchronization strategies.

  • Develop inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving or disaggregation.

  • Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads.

  • Work closely with Arago's hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features based on real model workloads.

Pay and benefits
  • Competitive cash compensation, with final package based on location, experience, and the pay of team members in similar positions.

  • Meaningful stock option plan offered (included in the majority of full time offers).

  • Healthcare coverage (including family-friendly options), pension contributions, professional development support, and 25 days of PTO, in addition to public holidays.

  • Ownership of a key technical domain, with significant vertical and/or horizontal growth opportunities, based on performance and individual drive.

We look forward to hearing how you can help shape the future of AI at Arago.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 90 000 - 130 000
Stock options
Healthcare coverage
Pension contributions
+2
Senior AI/ML Scientist
Senior AI/ML Scientist

Arago • Paris

Sur place
EUR 80 000 - 110 000
Competitive cash compensation
Stock options
Relocation bonus
+4
Senior AI/ML Scientist
Senior AI/ML Scientist

Arago Inc. • Paris

Sur place
EUR 70 000 - 100 000
Competitive cash compensation
Stock option plan
Healthcare coverage
+2
Senior Hardware Design Engineer
Senior Hardware Design Engineer

Arago Inc. • Paris

Sur place
EUR 70 000 - 120 000
Competitive cash compensation
Stock option plan
Relocation bonus
+4
Senior ASIC Architect
Senior ASIC Architect

Arago • Paris

Sur place
EUR 140 000 - 210 000
Stock option plan
Relocation support
Healthcare coverage
+1
Senior ASIC Architect
Senior ASIC Architect

arago • Paris

Sur place
EUR 60 000 - 80 000
Healthcare coverage
25 days of PTO
Relocation bonus
Software Engineer — Drivers & Virtualization
Software Engineer — Drivers & Virtualization

Arago • Paris

Sur place
EUR 90 000 - 130 000
Competitive cash compensation
Stock option plan
Relocation bonus
+4
Full-Stack Engineer, Office of the CEO
Full-Stack Engineer, Office of the CEO

Arago • Paris

Sur place
EUR 90 000 - 140 000
Stock options
Relocation bonus
Healthcare coverage
+3
System Administrator and Infrastructure Engineer
System Administrator and Infrastructure Engineer

Arago • Paris

Sur place
EUR 40 000 - 60 000
Competitive cash compensation
Stock option plan
Relocation bonus
+4