Multimodal AI Model Optimization Research Engineer

Tavus

San Francisco (CA)

Hybrid

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work schedules
Unlimited PTO
Healthcare
Gear stipend

Job summary

Tavus in San Francisco is seeking an experienced Research Scientist/Engineer focused on model optimization to join our core AI team. You will transform cutting-edge research models into fast, production-ready systems.

The role thrives in a startup environment, valuing independent thinking and calculated risk-taking. You will collaborate closely with researchers and engineers to turn new ideas into deployable solutions while optimizing latency, cost, and quality.

Qualifications

  • Strong experience in deep learning with PyTorch.
  • Hands-on model optimization and compression: distillation, pruning, quantization, mixed precision.
  • Understanding efficient architectures like low-rank adapters.
  • Excellent inference performance knowledge and GPU/accelerator fundamentals.
  • Strong Python coding and reliable research engineering practices.
  • Experience with large models and cloud environments.
  • Ability to read ML papers, reproduce results, and adapt ideas.
  • Clear communication and collaboration skills.

Responsibilities

  • Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization.
  • Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality.
  • Partner closely with researchers and engineers to turn new ideas into deployable systems

Skills

PyTorch
Model optimization
Distillation
Pruning / Sparsification
Quantization
Mixed precision
Low-rank adapters
Inference performance
Python coding
Cloud environments
Reading papers
Collaboration

Tools

TensorRT
ONNX Runtime
TVM
Triton
XLA
CUDA kernels

Job description

About Us

At Tavus, we're building the human layer of AI. Our mission is to make human-AI interaction as natural as face-to-face interaction, enabling the human touch where it has been previously unscalable.

We achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar behavior. Our models power everything from text-to-video AI avatars to real-time conversational video experiences across industries like healthcare, recruiting, sales, and education.

By enabling AI to see, hear, and communicate with human-like authenticity, we're creating the foundation for the next generation of AI employees, assistants, and companions.

We are a Series B company backed by top investors, including Sequoia, Y Combinator, and Scale VC. Join us in driving the future of human-AI interaction.

The Role

We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team.

Our ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks. We’re moving fast and looking for people who can help pave the path.

Your Mission
  • Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization

  • Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality

  • Partner closely with researchers and engineers to turn new ideas into deployable systems

Requirements
  • Strong experience in deep learning using PyTorch

  • Hands‑on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision

  • Understanding of efficient architectures such as low‑rank adapters

  • Strong understanding of inference performance and GPU/accelerator fundamentals

  • Strong Python coding skills and reliable research engineering practices

  • Experience working with large models and datasets in cloud environments

  • Ability to read ML papers, reproduce results, and adapt ideas

  • Clear communication and collaboration skills

Preferred Experience
  • Optimization of diffusion models, video/audio generative models, or large language models

  • Experience with real‑time or streaming systems (low‑latency APIs, WebRTC, streaming TTS/video)

  • Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA

  • Experience writing custom Triton/CUDA kernels or low‑level performance tuning

  • Experience with experiment tracking, benchmarking, and profiling at scale

  • Prior experience in research engineering or applied science roles

Location

This position is preferably hybrid in San Francisco, with relocation support offered. Remote candidates are also considered.

Benefits

When you join Tavus, you’re joining a family. We offer flexible work schedules, unlimited PTO, competitive healthcare and gear stipends, and a collaborative environment focused on learning and impact.

Culture & Diversity

We are not looking for cultural fits — we are looking for culture creators. Diversity drives our success, and we combine varied backgrounds, skills, and perspectives to build the best experiences for our clients.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote-Ready Multimodal AI Optimization Engineer
Remote-Ready Multimodal AI Optimization Engineer

Tavus • San Francisco (CA)

Hybrid
USD 180,000 - 280,000
Flexible work schedules
Unlimited PTO
Healthcare
+1
Conversational Modelling Research Engineer
Conversational Modelling Research Engineer

Tavus • San Francisco (CA)

Hybrid
USD 120,000 - 230,000
Senior Research Scientist
Senior Research Scientist

adaption • United States

Hybrid
USD 130,000 - 150,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Senior Research Scientist
Senior Research Scientist

Adaption Labs • United States

Hybrid
USD 100,000 - 140,000
Annual travel stipend
Weekly meal allowance
Comprehensive medical benefits
+1
Research Engineer, Developer Experience, Tinker
Research Engineer, Developer Experience, Tinker

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research Scientist
Research Scientist

Adaption Labs • New York (NY)

Remote
USD 120,000 - 160,000
Annual travel stipend to explore a new country
Weekly meal allowance
Comprehensive medical benefits
+1
Research, Audio Expertise
Research, Audio Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Research Vision Expertise
Research Vision Expertise

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Vision Expertise
Research, Vision Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1