Multimodal AI Model Optimization Research Engineer

Tavus

United States

Remote

USD 140,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health coverage
Unlimited PTO
Flexible hours
Remote-friendly
Learning stipend
Company retreats
Gear stipend

Job summary

Tavus is seeking an experienced Research Scientist/Engineer focused on model optimization to join our core AI team. You will take cutting-edge models and accelerate them for production using sparsification, distillation, and quantization.

You will own the optimization lifecycle, define metrics, run experiments, and benchmark latency, cost, and quality. Collaboration with researchers and engineers will turn ideas into deployable systems at Tavus.

Qualifications

  • Thrives in startup environments and takes calculated risks.
  • Ability to read ML papers, reproduce results, and adapt ideas.
  • Strong Python coding skills and reliable research engineering practices.
  • Experience with large models and datasets in cloud environments.
  • Hands-on experience with model optimization and compression: distillation, pruning, quantization, mixed precision.
  • Understanding of efficient architectures such as low-rank adapters.
  • Experience with deep learning using PyTorch.

Responsibilities

  • Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality.
  • Partner closely with researchers and engineers to turn new ideas into deployable systems.
  • Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization.
  • We’re looking for someone who can move fast and pave the path in a startup.

Skills

Python
PyTorch
Deep learning
Inference optimization
Experiment tracking
Communication
Cloud environments
Diffusion models

Tools

TensorRT
ONNX Runtime
TVM
Triton
CUDA

Job description

  • We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team
  • Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization
  • Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality
  • Partner closely with researchers and engineers to turn new ideas into deployable systems
Benefits
  • Comprehensive medical, dental and vision coverage for employees and their families, with 100% of monthly premiums paid by Tavus
  • Unlimited paid time off, as well as parental and pawental leave
  • Flexible hours
  • Remote-friendly, with teams at home or in co-working spaces across the globe
  • Debate and feedback are cornerstones of our culture
  • Yearly stipend for you to spend on any learning materials (i.e. newsletters, online courses, etc.)
  • Company retreats
  • Generous gear stipend
  • Frequent in person meet ups
  • Team members from diverse backgrounds – we’re looking for culture creators, not culture fits
  • Progressive, open-minded meritocracy
Qualifications
  • Our ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks
  • We’re moving fast and looking for people who can help pave the path
  • Ability to read ML papers, reproduce results, and adapt ideas
  • Clear communication and collaboration skills
  • Strong understanding of inference performance and GPU/accelerator fundamentals
  • Strong Python coding skills and reliable research engineering practices
  • Experience working with large models and datasets in cloud environments
  • Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision
  • Understanding of efficient architectures such as low-rank adapters
  • Strong experience in deep learning using PyTorch
  • Optimization of diffusion models, video/audio generative models, or large language models
  • Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)
  • Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA
  • Experience writing custom Triton/CUDA kernels or low-level performance tuning
  • Experience with experiment tracking, benchmarking, and profiling at scale
  • Prior experience in research engineering or applied science roles
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Multimodal AI Optimization Research Engineer
Multimodal AI Optimization Research Engineer

Tavus • United States

Remote
USD 140,000 - 210,000
Health coverage
Unlimited PTO
Flexible hours
+4
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Senior Research Scientist
Senior Research Scientist

adaption • United States

On-site
USD 130,000 - 150,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Research Engineer/Scientist(all levels), Efficient Models
Research Engineer/Scientist(all levels), Efficient Models

TikTok • San Jose (CA)

On-site
USD 150,000 - 210,000
Research Member of Technical Staff- Efficient Modeling
Research Member of Technical Staff- Efficient Modeling

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
Model Optimization Engineer
Model Optimization Engineer

Bright Vision Technologies • United States

Remote
USD 150,000 - 175,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Research Engineer - Inference
Research Engineer - Inference

ElevenLabs • Maine

On-site
USD 140,000 - 190,000
Annual discretionary stipend
Annual company offsite
Co-working stipend