Multimodal ML Engineer

Moonfire

Greater London

Hybrid

GBP 120,000 - 170,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Flexible time off
Paid time off
Relocation package

Job summary

White Circle is an AI Safety company building the safety, reliability, and optimization layer for AI systems. The team trains and deploys multimodal models across vision, audio, and video modalities to power safety policies at scale.

We seek an experienced Multimodal ML Engineer to lead end-to-end model development, from training from scratch to production serving, and to design robust evaluation benchmarks for real-world AI safety use cases.

Qualifications

  • Proficient in PyTorch with experience in distributed training.
  • Strong background in multimodal model architectures across vision/audio domains.
  • Hands-on experience with RLHF/alignment for multimodal systems.
  • Track record shipping models to production with latency targets.

Responsibilities

  • Train and fine-tune large-scale multimodal models (vision-language, audio, speech).
  • Extend models across modalities including video temporal modeling and streaming audio.
  • Design experiments and data mixes; build training recipes.
  • Build and maintain multimodal data pipelines and synthetic data generation.
  • Train and optimize MoE architectures for efficient inference.
  • Deploy models end-to-end from research checkpoint to production serving.
  • Define metrics and benchmarks relevant to product needs.

Skills

PyTorch
Distributed training
Multimodal architectures
RLHF/alignment
Video/audio sequence modeling
Production deployment
Engineering fundamentals
Latency optimization

Tools

DeepSpeed
FSDP
Whisper
HuBERT

Job description

TLDR:

Multimodal ML Engineer to train and ship vision, audio, video, and speech models for an AI safety platform that operates at 100M+ API calls/month.

About us

White Circle is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies - simple natural-language rules that define what an AI model should and shouldn't do. We automatically test, enforce, and continuously improve these policies at scale.

  • We've recently raised our Series A funding round, taking our total funding to $70M. Our investors include top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others
  • We process over 100M+ API calls every month
  • We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model

We're a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built - you're the one we need.

You will
  • Train and fine-tune large-scale multimodal models (vision-language, audio, speech) from scratch and from pretrained checkpoints
  • Extend models across modalities: image understanding, video temporal modeling, long-context processing, and streaming audio
  • Design and run experiments: architecture changes, data mixes, training recipes
  • Build and maintain multimodal data pipelines - from raw images, video, and audio recordings to training-ready datasets, including synthetic data generation
  • Train and optimize MoE architectures for efficient multimodal inference
  • Build alignment pipelines: SFT, DPO, GRPO, reward modeling - across modalities, not just text
  • Optimize models for production: quantization, distillation, batching, streaming and low-latency serving
  • Deploy models end-to-end: from research checkpoint to production serving
  • Define evaluation metrics and benchmarks that actually matter for the product: visual QA, spatial reasoning, video comprehension, speech and audio understanding
You’ll fit right in if you
  • 3+ years training large-scale deep learning models in multimodal domains (vision-language, audio, speech, or acoustic)
  • Strong PyTorch skills with hands-on distributed training experience (DeepSpeed, FSDP, or similar)
  • Deep experience with multimodal architectures - you understand how vision/audio encoders, projectors, and LLMs fit together (LLaVA, Qwen-VL, InternVL, Audio Flamingo, Omni Qwen, Audio Qwen, Whisper, HuBERT, Conformer, or similar)
  • Hands-on with RLHF/alignment for multimodal: GRPO, DPO, reward modeling - not just for text
  • Experience with video and/or audio sequence modeling: temporal modeling, long-context processing, efficient attention, streaming inference
  • Track record of shipping models to production: you've hit latency targets and optimized inference, not just reported benchmark scores
  • Comfortable with large-scale multimodal dataset curation: image-text pairs, video-instruction data, audio preprocessing, augmentation, synthetic data generation
  • Familiar with MoE architectures and their tradeoffs for multimodal workloads
  • Strong engineering fundamentals: clean code, version control, testing, documentation
A big plus
  • Understanding of audio signal processing fundamentals (spectrograms, mel features, noise reduction)
Why White Circle
  • Competitive compensation package, including equity
  • Flexible time off
  • Paid time off in line with your local regulations, no matter where you work from
  • Work from Paris or London (hybrid) + relocation package for Paris
  • Best medical insurance in France
  • Learning and development support for courses, conferences, and opportunities to grow your skills
  • All the hardware, tools, and services you need
  • Covered subscriptions for AI agents and IDEs
  • Team off-sites twice a year: we've recently been to the Alps and to Saint-Tropez
Process
  1. Introductory call with HR (25 min)
  2. Take-home test assignment
  3. Technical interview with Head of Applied Research (60 min)
  4. Final conversation with CEO (45 min)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Research Engineer
ML Research Engineer

Moonfire • Greater London

Hybrid
GBP 90,000 - 130,000
Hybrid office in London/Paris
Relocation support
Health insurance
+2
Research Engineer (Evals)
Research Engineer (Evals)

Moonfire • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Hybrid work model
Office in London/Paris
+3
Research Scientist/Engineer (Agentic Systems)
Research Scientist/Engineer (Agentic Systems)

White Circle • Greater London

On-site
GBP 112,141 - 186,902
Relocation package
Medical insurance (France)
All hardware and tools provided
+2
QA Engineer (Manual + Automation)
QA Engineer (Manual + Automation)

Moonfire • Greater London

Hybrid
GBP 50,000 - 70,000
Equity
Flexible time off
Relocation package for Paris
+1
Senior Data Labeler
Senior Data Labeler

Moonfire • Greater London

Hybrid
GBP 60,000 - 100,000
Equity
Flexible time off
Office in Paris
+8
Recruiter (GTM & Business)
Recruiter (GTM & Business)

White Circle • Greater London

On-site
GBP 60,000 - 90,000
Meaningful equity package
Paid time off per local regulations
Hardware and software tools for work
+2
Talent Lead
Talent Lead

White Circle • Greater London

On-site
GBP 90,000 - 130,000
Equity
Flexible time off
Language lessons
+3
AI engineer
AI engineer

Lucis (YC P25) • Greater London

On-site
GBP 70,000 - 90,000
AI engineer
AI engineer

Lucis • Greater London

On-site
GBP 120,000 - 160,000
Relocation support
Impactful work ownership
Machine Learning Engineer / ML Engineer - Roleplay Sessions
Machine Learning Engineer / ML Engineer - Roleplay Sessions

Synthesia • Greater London

On-site
GBP 110,000 - 140,000
RSUs
25 days leave