Ai Distilation Expert

RB Labs

San Francisco (CA)

On-site

USD 130,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

RB Labs is building a real-time generative AI system where latency is the core challenge. We’re seeking a model distillation specialist to shrink large models into fast, deployment-ready versions without sacrificing quality.

This is a hands-on role: you’ll design and run distillation experiments, evaluate quality/latency tradeoffs, and ship the results to production GPU infrastructure.

Qualifications

  • 2 years hands-on PyTorch experience including training/fine-tuning from scratch.
  • Proven experience with model distillation and latency optimization.
  • Strong knowledge of model compression techniques: quantization, pruning, knowledge distillation.
  • Ability to read and extend dense research code.

Responsibilities

  • Design and run knowledge distillation pipelines – teacher/student setups for large models.
  • Optimize for latency via distillation, quantization, pruning, and mixed precision.
  • Evaluate rigorously with quality-vs-latency comparisons and A/B tests.
  • Own training loops – write and debug full PyTorch training code, not just pre-existing scripts.
  • Ship distilled models to production GPUs and measure real-world performance.

Skills

PyTorch
Model distillation
Quantization
Pruning
LoRA/PE methods

Tools

PyTorch

Job description

About the Role

We’re building a real-time generative AI system where latency is the core challenge. We’re looking for a model distillation specialist to help us shrink large models into fast, deployment-ready versions without sacrificing quality.

This is a hands-on role: you’ll design and run distillation experiments, evaluate quality/latency tradeoffs, and ship the results to production GPU infrastructure.

What You’ll Do
  • Design and run knowledge distillation pipelines – teacher/student setups for large neural models (LLMs and/or generative vision models)
  • Optimize for latency – reduce inference cost via distillation, quantization, pruning, and mixed precision
  • Evaluate rigorously – build empirical quality-vs-latency comparisons and A/B test variants
  • Own training loops – write and debug full PyTorch training code, not just run existing scripts
  • Ship to production – deploy distilled models on live GPU servers and measure real-world performance
Must Have
  • 2 years hands-on PyTorch, including training/fine-tuning models from scratch – not just serving pretrained ones
  • Proven experience with model distillation – you’ve actually distilled a model and measured the results
  • Strong grasp of model compression techniques: quantization, pruning, knowledge distillation, LoRA/parameter-efficient methods
  • Solid understanding of loss design for distillation (KL divergence, feature/logit matching, perceptual losses)
  • Comfortable reading and extending dense, lightly-documented research code
Strong Plus
  • Experience distilling autoregressive LLMs or speech models
  • Experience with GANs, diffusion, or flow-matching models
  • GPU inference optimization: torch.compile, CUDA memory management, multi-GPU setups
  • Familiarity with streaming/real-time ML systems
Working Style

We operate like a research lab: empirical, fast iteration, incomplete docs, self-directed testing. You’ll need to be comfortable designing your own experiments and pushing for more engineering structure as you go.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Real-Time AI Distillation Engineer
Real-Time AI Distillation Engineer

RB Labs • San Francisco (CA)

On-site
USD 130,000 - 180,000
Efficient GenAI Research Engineer: Distillation
Efficient GenAI Research Engineer: Distillation

ByteDance • San Jose (CA)

On-site
USD 254,000 - 588,000
Remote AI Engineer: From-Scratch ML & LLM Training
Remote AI Engineer: From-Scratch ML & LLM Training

Saguna Consulting Services • United States

Remote
USD 150,000 - 230,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Applied AI Researcher, Post-Training
Applied AI Researcher, Post-Training

Distyl • New York (NY), San Francisco (CA)

Hybrid
USD 150,000 - 250,000
Equity
Medical insurance
Flexible time off
+6
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff, Training
Member of Technical Staff, Training

Inception • San Francisco (CA)

On-site
USD 180,000 - 230,000
Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Distillation Architect — AI Model Efficiency Lead
Distillation Architect — AI Model Efficiency Lead

Waabi • United States

On-site
USD 195,000 - 286,000
Competitive compensation and equity
Health and Wellness benefits (Medical,
Unlimited Vacation
+3
Machine Learning Researcher
Machine Learning Researcher

SOLANA FOUNDATION • San Francisco (CA)

On-site
USD 250,000 - 350,000
Equity in a high-growth startup
Comprehensive benefits