Intern - ML Inference Performance Engineer

Axelera AI

Eindhoven

Hybrid

EUR 42,000 - 64,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Pension plan
Employee insurances
Company shares

Job summary

Axelera AI is seeking a curious, rigorous engineer to join our team and dig into the performance of ML inference systems. You’ll work across the full inference stack—model export, compiler toolchains, and runtime execution on silicon—to build a clear, evidence-based picture of how different platforms perform and why.

Your work will go beyond running benchmarks: you’ll develop a repeatable evaluation methodology, investigate bottlenecks at hardware and software levels, and build tooling that

Qualifications

  • Currently enrolled in final years of Bachelor's or Master's programme or Masters thesis project.
  • Proficiency in Python and C/C++ development.
  • Experience with end-to-end computer vision pipelines.
  • Familiarity with benchmarking concepts (performance, latency).
  • Experience with inference tools, APIs, or SDKs (e.g., TensorRT).
  • Understanding of deep learning model concepts (quantization, ONNX, PyTorch).
  • Linux, Bash scripting, and Docker proficiency.
  • Hands-on experience with embedded hosts.
  • Strong written and verbal English communication.
  • Good organisational skills.

Responsibilities

  • Benchmarking & tooling: develop benchmarking tools for throughput, latency, power, and accuracy across platforms; maintain dashboards for visualisations.
  • Platform evaluation: evaluate AI accelerators and vendor SDKs and track model support across platforms.
  • Pipeline analysis: characterize full inference pipelines and ensure consistent configurations across platforms.
  • Lab & infrastructure: set up lab hosts across hardware platforms and onboard new evaluation hardware.
  • Reporting: synthesize findings into clear reports informing engineering and roadmap decisions.

Skills

Python
C/C++
CV pipelines
Benchmarking
Inference SDKs
DL concepts
Agentic AI
Git
LLM benchmarking
Linux
Bash scripting
Embedded hardware
English communication
Organisational skills

Education

Bachelor's or Master's in Computer Engineering/Computer Science/EE

Tools

TensorRT
GStreamer
Docker
Git

Job description

About Us

Axelera AI is not your regular deep-tech company. We are creating the next-generation AI platform to support anyone who wants to help advancing humanity and improve the world around us.

In just five years, we have raised a total of $370 million and have built a world-class team of 250+ employees (including 60+ PhDs with more than 40,000 citations), both remotely from 20 different countries and with offices in Belgium, France, Switzerland, Italy, the UK, headquartered at the High Tech Campus in Eindhoven, Netherlands.

We have also launched our Metis AI Platform, which achieves a 3-5x increase in efficiency and performance, and have visibility into a strong business pipeline exceeding $100 million.

Our unwavering commitment to innovation has firmly established us as a global industry pioneer.

Are you up for the challenge?

Position Overview

We're looking for a curious, rigorous engineer to join our team and dig into the performance of ML inference systems. You'll work across the full inference stack — from model export and compiler toolchains to runtime execution on silicon — to build a clear, evidence-based picture of how different platforms perform and why.

Your work will go beyond running benchmarks: you'll develop a repeatable evaluation methodology, investigate performance bottlenecks at the hardware and software level, and build the tooling that transforms raw measurements into actionable insight. The findings you produce can directly shape our product decisions.

Key responsibilities:
  • Benchmarking & Tooling: Develop a thorough understanding of internal benchmarking tools covering throughput, latency, power, and accuracy across device-level, host-transaction, and end-to-end pipeline scenarios. Improve existing tooling, define reproducible procedures, and establish a standardised results format for rigorous cross-platform comparisons. Maintain a dedicated dashboard for performance visualisations.

  • Platform Evaluation: Research and evaluate AI accelerator products from various vendors, gaining hands-on experience with their SDKs, toolchains, flexibility, and limitations through a structured evaluation process. Track model support across platforms to identify strengths, gaps, and areas for improvement.

  • Pipeline Analysis: Characterise full inference pipelines, capturing host-device transaction overhead and end-to-end performance metrics. Ensure equivalent pipeline configurations across platforms using frameworks such as GStreamer to maintain methodological consistency.

  • Lab & Infrastructure: Set up and maintain lab hosts across multiple hardware platforms and support the onboarding of new evaluation hardware.

  • Reporting: Synthesise findings into clear, structured reports that directly inform engineering and roadmap decisions.

Requirements:
  • Currently enrolled in the final years of a Bachelor's programme or in a Master's programme in Computer Engineering, Electrical Engineering, Computer Science, or a related field. This position may also be carried out as a Master's thesis project.

  • Python development experience

  • C/C++ knowledge

  • Experience with end-to-end computer vision pipelines

  • Familiarity with benchmarking concepts (performance, latency, etc.)

  • Experience with inference tools, APIs, or SDKs (e.g., TensorRT)

  • Familiarity with deep learning model concepts (quantization, ONNX, PyTorch, etc.)

  • Development experience using agentic AI

  • Knowledge of version control (Git)

  • Familiarity with LLM benchmarking concepts

  • Proficiency with Linux, Bash scripting, and Docker

  • Hands-on experience with embedded hosts

  • Proficient written and verbal communication skills in English, with the ability to document findings clearly and precisely.

  • Good organisational skills

Nice to have:
  • GStreamer knowledge

  • Basic GUI design experience

Location

Work from our Axelera AI office in Eindhoven (Netherlands).

What weoffer

This is your chance to shape and be part of a dynamic, fast-growing, international organization. We offer an attractive compensation package, including a pension plan, extensive employee insurances and the option to get company shares.

An open culture that supports creativity and continual innovation is awaiting you. Collaborative ownership and freedom with responsibility is characteristic for the way we act and work as a team.

At Axelera AI, we wholeheartedly embrace equal opportunity and hold diversity in the highest regard. Our steadfast commitment is to cultivate a warm and inclusive environment that empowers and celebrates every member of our team. We welcome applicants from all backgrounds to join us in shaping the future of AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Intern - ML Inference Performance Engineer
Intern - ML Inference Performance Engineer

Axelera • Eindhoven

Hybrid
EUR 42,000 - 62,000
Pension plan
Employee share options
Insurance package
Director - Hardware Engineering, AI Infrastructure Systems
Director - Hardware Engineering, AI Infrastructure Systems

Axelera AI • Netherlands

Hybrid
EUR 190,000 - 270,000
Pension plan
Employee insurances
Company shares
Director - Hardware Engineering, AI Infrastructure Systems
Director - Hardware Engineering, AI Infrastructure Systems

Axelera AI • Eindhoven, Amsterdam

Remote
EUR 180,000 - 260,000
Pension plan
Extensive employee insurances
Equity / stock options
+1
Director, AI Infrastructure Systems — Remote-Ready Leadership
Director, AI Infrastructure Systems — Remote-Ready Leadership

Axelera AI • Eindhoven, Amsterdam

On-site
EUR 180,000 - 260,000
ML Inference Performance Intern — Benchmarking & Tools
ML Inference Performance Intern — Benchmarking & Tools

Axelera AI • Eindhoven

Hybrid
EUR 42,000 - 64,000
Pension plan
Employee insurances
Company shares
Automotive Application Engineer
Automotive Application Engineer

Axelera AI • Amsterdam, Eindhoven

Remote
EUR 70,000 - 110,000
Pension plan
Extensive employee insurances
Stock options
+1
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

United States Digital Space LLC • Amsterdam

Hybrid
EUR 100,000 - 180,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Developer Relations and Community Manager
Developer Relations and Community Manager

Axelera AI • Eindhoven

Remote
GBP 60,000 - 90,000
Pension plan
Employee insurances
Stock options
ML Performance Engineer
ML Performance Engineer

Internetwork Expert • Amsterdam

Hybrid
EUR 100,000 - 150,000
High base salary
Generous bonus structure
Cutting-edge hardware and software
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Netherlands

On-site
EUR 120,000 - 180,000
Competitive compensation
Learning opportunities
Ownership in work
+1