ML Performance Engineer

Internetwork Expert

Amsterdam

Hybrid

EUR 100,000 - 150,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

High base salary
Generous bonus structure
Cutting-edge hardware and software

Job summary

Internetwork Expert is seeking an ML Engineer to speed up large-scale model training by optimizing the internal stack and compute infrastructure. You will work across the full training pipeline—from GPU kernels to system-level throughput—using profiling, CUDA-level tuning, and distributed systems techniques.

This role focuses on reducing training time, boosting iteration speed, and improving compute efficiency within a growing ML training systems team in Amsterdam.

Qualifications

  • Experience optimizing neural network training in large-scale production or research settings.
  • Hands-on experience with PyTorch or JAX.
  • Experience with CUDA, Triton, or other low-level GPU technologies for performance tuning.
  • Strong Python for infrastructure tooling and integration with ML frameworks.

Responsibilities

  • Optimize the model training pipeline for speed and reliability across the full stack.
  • Apply GPU-level optimizations to improve training performance at scale.
  • Identify and resolve performance bottlenecks from data loading to CUDA kernels.
  • Build tools and extend internal infrastructure for scalable, reproducible training workflows.
  • Mentor engineers and researchers in performance best practices.
  • Grow the team’s GPU and systems-level capabilities and drive fast experimentation.

Skills

Python
PyTorch/JAX
Profiling/Debugging
Distributed training concepts

Tools

Nsight
CUDA profiling tools
Torch profiler

Job description

We're looking for a performance-focusedML Engineerto help speed up large-scale model training by optimizing our internal stack and compute infrastructure. You'll work across the full training pipeline - from GPU kernels to system-level throughput - applying profiling, CUDA-level tuning, and distributed systems techniques. The goal is to reduce training time, boost iteration speed, and use compute more efficiently.

This is a key role in a growing team building deep technical expertise in ML training systems.

Responsibilities
  • Optimize our model training pipeline to improve both speed and reliability, enabling faster and more efficient experimentation;
  • Apply GPU-level optimization techniques using tools like JAX, Triton, low-level CUDA to improve training performance and efficiency at scale;
  • Identify and resolve performance bottlenecks across the entire ML pipeline - from data loading and preprocessing to CUDA kernels;
  • Build tools and extend internal infrastructure to support scalable, reproducible, and high-performance training workflows;
  • Mentor and support engineers and researchers in adopting performance best practices across the team;
  • Help grow the team’s GPU and systems-level capabilities, and contribute to a culture of engineering excellence and rapid experimentation.
Requirements
  • Demonstrated experience optimizing neural network training in production or large-scale research settings - e.g. reducing training time, improving hardware utilization, or accelerating feedback cycles for ML researchers;
  • Extensive practical experience with ML frameworks such as PyTorch or JAX;
  • Hands-on experience with training and optimizing deep learning architectures such as LSTM and Transformer-based models, including different attention mechanisms;
  • Experience working with CUDA, Triton, or other low-level GPU technologies for performance tuning;
  • Proficiency in profiling and debugging training pipelines, using tools such as Nsight/cprofiler/CUDA/gdb/torch profiler;
  • Understanding of distributed training concepts (e.g. data/model/tensor/sequence/pipeline/context parallelism, memory and compute tradeoffs);
  • A collaborative and proactive mindset, with strong communication skills and the ability to mentor teammates and partner effectively within the team;
  • Strong proficiency in Python for building infrastructure-level tooling, debugging training systems, and integrating with ML frameworks and profiling tools;
What we offer
  • High base salary and social benefits;
  • Generous bonus structure. We are very flexible in discussing salary and conditions of employment;
  • Cutting-edge hardware and software in production as well as high technical expertise of the company which allows implementation of bold ideas and boosting great results. Ownership over initiatives that directly solve business problems;
  • Ability to trade on dozens of international exchanges;
  • Flexible workflow (lack of formalism and bureaucracy, no pressure and over-management) and working schedule;
  • Tuition reimbursement, conference and training sponsorship;
  • Location: Amsterdam preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Training Systems Engineer - Performance Focus
ML Training Systems Engineer - Performance Focus

Internetwork Expert • Amsterdam

Hybrid
EUR 100,000 - 150,000
High base salary
Generous bonus structure
Cutting-edge hardware and software
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

United States Digital Space LLC • Amsterdam

Hybrid
EUR 100,000 - 180,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Performance Engineer
Performance Engineer

Stream HPC BV • Amsterdam

Hybrid
EUR 65,000 - 95,000
MLOps Engineer
MLOps Engineer

The French Sourcer • Amsterdam

Hybrid
EUR 90,000 - 130,000
Health insurance
Pension plan
Equity plan
+1
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

Slashhash • Amsterdam

Hybrid
EUR 120,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

IMC Trading • Amsterdam

On-site
EUR 70,000 - 90,000
Senior ML Solutions Architect - Token Factory
Senior ML Solutions Architect - Token Factory

Jobgether • Netherlands

Remote
EUR 120,000 - 190,000
Fully remote work from Europe
Career growth and learning
Autonomy and ownership
+2
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Harnham • Amsterdam

On-site
EUR 100,000 - 130,000
Competitive salary with performance-related bonuses
Opportunity to work in a highly innovative environment
Flexible working arrangements
MLOps Engineer
MLOps Engineer

EPAM Systems • Netherlands

Hybrid
EUR 85,000 - 120,000
Sr. Machine Learning Engineer | Up to €170K
Sr. Machine Learning Engineer | Up to €170K

Bluebird • Amsterdam

On-site
EUR 144,000 - 170,000