Real-Time AI Distillation Engineer

RB Labs

San Francisco (CA)

On-site

USD 130,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

RB Labs is building a real-time generative AI system where latency is the core challenge. We’re seeking a model distillation specialist to shrink large models into fast, deployment-ready versions without sacrificing quality.

This is a hands-on role: you’ll design and run distillation experiments, evaluate quality/latency tradeoffs, and ship the results to production GPU infrastructure.

Qualifications

  • 2 years hands-on PyTorch experience including training/fine-tuning from scratch.
  • Proven experience with model distillation and latency optimization.
  • Strong knowledge of model compression techniques: quantization, pruning, knowledge distillation.
  • Ability to read and extend dense research code.

Responsibilities

  • Design and run knowledge distillation pipelines – teacher/student setups for large models.
  • Optimize for latency via distillation, quantization, pruning, and mixed precision.
  • Evaluate rigorously with quality-vs-latency comparisons and A/B tests.
  • Own training loops – write and debug full PyTorch training code, not just pre-existing scripts.
  • Ship distilled models to production GPUs and measure real-world performance.

Skills

PyTorch
Model distillation
Quantization
Pruning
LoRA/PE methods

Tools

PyTorch

Job description

RB Labs is building a real-time generative AI system where latency is the core challenge. We’re seeking a model distillation specialist to shrink large models into fast, deployment-ready versions without sacrificing quality.

This is a hands-on role: you’ll design and run distillation experiments, evaluate quality/latency tradeoffs, and ship the results to production GPU infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Ai Distilation Expert
Ai Distilation Expert

RB Labs • San Francisco (CA)

On-site
USD 130,000 - 180,000
Efficient GenAI Research Engineer: Distillation
Efficient GenAI Research Engineer: Distillation

ByteDance • San Jose (CA)

On-site
USD 254,000 - 588,000
GenAI Research Engineer: Distillation & Efficient Models
GenAI Research Engineer: Distillation & Efficient Models

ByteDance • Seattle (WA)

On-site
USD 242,000 - 456,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6
Generative AI Research Engineer — Efficient Models
Generative AI Research Engineer — Efficient Models

ByteDance • San Jose (CA)

On-site
USD 162,000 - 388,000
Distillation Architect — AI Model Efficiency Lead
Distillation Architect — AI Model Efficiency Lead

Waabi • United States

On-site
USD 195,000 - 286,000
Competitive compensation and equity
Health and Wellness benefits (Medical,
Unlimited Vacation
+3
Remote Research Engineer: Real-Time AI Inference
Remote Research Engineer: Real-Time AI Inference

ElevenLabs • Maine

Hybrid
USD 140,000 - 190,000
Annual discretionary stipend
Annual company offsite
Co-working stipend
AI Model Distillation & Efficiency Lead
AI Model Distillation & Efficiency Lead

Waabi • San Francisco (CA), Northern (KY)

Hybrid
USD 195,000 - 286,000
Competitive compensation and equity
Health and Wellness benefits (Medical,
Unlimited Vacation
+4
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Remote Real-Time AI Inference Engineer
Remote Real-Time AI Inference Engineer

ElevenLabs • United States

On-site
USD 120,000 - 190,000
Generative AI Research Engineer
Generative AI Research Engineer

TikTok • San Jose (CA)

On-site
USD 150,000 - 210,000