ML Researcher

Desert Ant Labs

Amsterdam

Remote

EUR 70,000 - 120,000

Full time

44 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Desert Ant Labs in Amsterdam offers a hands-on role to advance speech, vision, video, and text models for on-device deployment. You will help shrink models to fit real devices and ship them in our Detail and Subwave apps.

The role covers data generation, benchmarking across languages, and collaborating with teams to define model capabilities and evaluation standards. Remote work options exist alongside Amsterdam-based collaboration.

Qualifications

  • Write solid Python and know PyTorch well.
  • Ship models and explain end-to-end what was done.
  • Write clean, tested code and review changes carefully.
  • Experience with on-device runtimes and edge ML.

Responsibilities

  • Generate synthetic and augmented data, test on real recordings, videos, and text from the devices people use.
  • Measure latency, memory, heat, and battery use on real hardware before reporting a number.
  • Build models that combine audio, video, and text, such as labeling who speaks when.
  • Decide with Detail and Subwave what a model should do, then measure the model in the app.
  • Write the model card and the release post for each model.

Skills

Python
PyTorch
Model benchmarking
Data generation

Tools

Core ML
LiteRT
ONNX Runtime
ExecuTorch
MLX
llama.cpp
coremltools
CUDA

Job description

Train our speech, vision, video, and text models, make them smaller and faster on real devices, and ship them in Detail and Subwave, our own apps.

  • Train speech, vision, video, and text models that run on a phone, and make them smaller and faster.
  • You've shipped a model under a size or latency budget, and go deep in speech, vision and video, text, or efficiency.

Location Amsterdam, Remote

Time zone UTC-5 to UTC+1

Type Full time

About Desert Ant Labs
About the role
Our models today
  • Text: Redact filters PII in 27 languages, Gist tags topics in 101 languages, Tongue detects 84 languages from three words, and Title suggests a title and description for any text. Models for structured extraction and hate speech detection are in beta.
  • Efficiency: every model is quantized, distilled, or pruned to fit, converted for each platform, and benchmarked on real devices. Shrinking a model from 8MB to 2MB can take weeks, and the work is wasted when the 8MB model already fits the product, so we work on the limits that matter first.
  • Next: getting the beta models out of beta, and new models in every area.
How we do research

We pick each model's default settings for the products that use the model, such as which kinds of personal data Redact removes when the developer changes nothing, or whether Uhm leaves a filler word in or risks cutting a real word. You make those decisions with the team, and we expect you to say so when you disagree.

We want a published benchmark for every language a model supports, and the same quality in each. We aren't there yet for every model, and closing those gaps is part of the work.

What you will do
  • Generate synthetic and augmented data, and test on real recordings, videos, and text from the devices and rooms people use.
  • Measure latency, memory, heat, and battery use on real hardware before you report a number.
  • Build models that combine audio, video, and text, such as labeling who speaks when.
  • Decide with the Detail and Subwave teams what a model should do, then measure the model in the app.
  • Write the model card and the release post for each model.
You might be a fit if you
  • Can design an evaluation and defend its results in public, including in languages other than English.
  • Write solid Python and know PyTorch well.
  • For a senior role: have shipped models for at least 5 years, and can walk through one, from what it was for to how you measured it and what you'd change.
  • Write clean, tested code, and spot a weak change in review, whether a person or an agent wrote it.
  • Plan and run your work through coding agents such as Claude Code, with a low tolerance for slop.
  • Worked on low-resource languages, or fine-tuned a small language model for one task.
  • Know an on-device runtime in depth: Core ML, LiteRT, ONNX Runtime, ExecuTorch, or MLX.
  • Wrote Metal, CUDA, or NPU kernels, or contributed to ExecuTorch, MLX, llama.cpp, or coremltools.
How we work

Start what needs starting without waiting to be asked, and finish what you start. Take on work outside your role when a project needs you.

We ship quickly, so we cut scope until only the part users notice is left. Anyone can comment on your work or redo your draft, and we say early when work isn't ready. We read that feedback as help.

We judge the work by what shipped and what changed because of it. Nobody counts hours.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Device Lab and Benchmark Engineer
Device Lab and Benchmark Engineer

Desert Ant Labs • Amsterdam

On-site
EUR 90,000 - 120,000
Developer Relations Engineer
Developer Relations Engineer

Desert Ant Labs • Amsterdam

Remote
EUR 60,000 - 90,000
Mobile Engineer: Android
Mobile Engineer: Android

Desert Ant Labs • Amsterdam

Remote
EUR 70,000 - 95,000
Runtime and Performance Engineer
Runtime and Performance Engineer

Desert Ant Labs • Amsterdam

Hybrid
EUR 90,000 - 140,000
Office in Amsterdam
Remote work option
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Desert Ant Labs • Amsterdam

Remote
EUR 80,000 - 120,000
Modelling Resident
Modelling Resident

adaption • Netherlands

On-site
EUR 18,000 - 32,000
On-Device AI Researcher — Fast, Tiny Models for Phones
On-Device AI Researcher — Fast, Tiny Models for Phones

Desert Ant Labs • Amsterdam

Remote
EUR 70,000 - 120,000
AI Researcher
AI Researcher

Reson8 • Amsterdam

Hybrid
EUR 90,000 - 150,000
Salary competitive with the top of the
Equity as an early employee
25 days vacation
Forward Deployed Engineer, Benelux
Forward Deployed Engineer, Benelux

Telnyx • Amsterdam

On-site
EUR 120,000 - 160,000
Senior Machine Learning Engineer – LLMs
Senior Machine Learning Engineer – LLMs

Prosus • Amsterdam

Hybrid
EUR 110,000 - 150,000
Hybrid work model (Amsterdam)
Competitive compensation
Top-spec MacBook Pro
+1