Runtime and Performance Engineer

Desert Ant Labs

Amsterdam

Hybrid

EUR 90,000 - 140,000

Full time

29 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Office in Amsterdam
Remote work option

Job summary

Desert Ant Labs seeks a developer to optimize models for Apple platforms, Android, and Windows, shaping the Swift core and platform bindings. You will push workloads to GPUs and Neural Engines, and design SDK APIs to maximize speed on diverse devices.

You will write Metal/WebGPU kernels when needed and collaborate with researchers to advance model architectures for real devices, contributing to a fast, robust SDK.

Qualifications

  • Experience speeding software on real hardware and profiling performance.
  • Shipped an SDK or runtime used by developers.
  • Able to read compiler output and profiler traces.
  • Write Swift with C interop and bindings to other languages.
  • Work in C, C++, Kotlin, or Rust where a platform needs it.
  • Knowledge of GPUs/ML accelerators and kernel programming.
  • Contributed to llama.cpp, MLX, or ONNX Runtime.

Responsibilities

  • Run as much of each model on the Neural Engine, NPU, or GPU as the hardware allows.
  • Build the shared Swift core and its bindings to Kotlin, WebAssembly, and native code: memory budgets, routing, and fallbacks.
  • Write Metal, WebGPU, or native kernels when a runtime runs slowly; you may write the first ones.
  • Collaborate with researchers on model architectures that run fast on real devices.

Skills

Swift
Kotlin
C++
Rust
WebAssembly

Tools

Metal
DirectX
WebGPU
ONNX Runtime
llama.cpp
MLX

Job description

Make our models run faster on Apple platforms, Android, and Windows, in our shared Swift core and its platform bindings.

  • You have deep experience on Apple platforms (Swift, Metal), Android (Kotlin, the NDK), or Windows (C++, DirectX).

Location Amsterdam, Remote

Time zone UTC-5 to UTC+1

Type Full time

About the role

You work from GPU kernels up to the SDK API. You also design the SDK API, so the default call runs a model the fastest way the device allows.

What you will do
  • Run as much of each model on the Neural Engine, NPU, or GPU as the hardware allows.
  • Build the shared Swift core and its bindings to Kotlin, WebAssembly, and native code: memory budgets, routing, and fallbacks.
  • Write Metal, WebGPU, or native kernels when a runtime runs a model too slowly. We have none yet, so you write the first.
  • Work with researchers on model architectures that run fast on real devices.
You might be a fit if you
  • Have made software measurably faster on real hardware, and can explain the profile and the fix.
  • Have shipped an SDK or runtime that other developers built on, and have changed an API after seeing developers use the API wrong.
  • Read compiler output and profiler traces.
  • Write Swift, including C interop and bindings to other languages.
  • Work in C, C++, Kotlin, or Rust where a platform needs it.
  • Write clean, tested code, and spot a weak change in review, whether a person or an agent wrote it.
  • Plan and run your work through coding agents such as Claude Code, with a low tolerance for slop.
  • Contributed to llama.cpp, MLX, or ONNX Runtime.
  • Wrote GPU compute code on mobile.
How we work

Start what needs starting without waiting to be asked, and finish what you start. Take on work outside your role when a project needs you.

We ship quickly, so we cut scope until only the part users notice is left. Anyone can comment on your work or redo your draft, and we say early when work isn't ready. We read that feedback as help.

We judge the work by what shipped and what changed because of it. Nobody counts hours.

What we offer
  • In our Amsterdam office, or remote anywhere between Eastern Time in North America (UTC-5) and Central European Time (UTC+1), so everyone shares part of the working day.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Mobile Engineer: Android
Mobile Engineer: Android

Desert Ant Labs • Amsterdam

Remote
EUR 70,000 - 95,000
Device Lab and Benchmark Engineer
Device Lab and Benchmark Engineer

Desert Ant Labs • Amsterdam

On-site
EUR 90,000 - 120,000
Fast-Track Runtime & Performance Engineer
Fast-Track Runtime & Performance Engineer

Desert Ant Labs • Amsterdam

Hybrid
EUR 90,000 - 140,000
Office in Amsterdam
Remote work option
ML Performance Engineer
ML Performance Engineer

Internetwork Expert • Amsterdam

On-site
EUR 100,000 - 150,000
High base salary
Generous bonus structure
Cutting-edge hardware and software
ML Researcher
ML Researcher

Desert Ant Labs • Amsterdam

Remote
EUR 70,000 - 120,000
Developer Relations Engineer
Developer Relations Engineer

Desert Ant Labs • Amsterdam

Remote
EUR 60,000 - 90,000
Performance Engineer
Performance Engineer

Stream HPC BV • Amsterdam

Hybrid
EUR 65,000 - 95,000
Senior Software Engineer - Serverless
Senior Software Engineer - Serverless

Runware • Netherlands

On-site
EUR 90,000 - 140,000
Generous paid time off
Stock options
Remote-first setup
+3
System Performance Architect
System Performance Architect

DeepRec.ai • Amsterdam

Hybrid
EUR 100,000 - 140,000
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

United States Digital Space LLC • Amsterdam

On-site
EUR 100,000 - 180,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3