Research Scientist

Velvet

San Francisco (CA)

On-site

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A data research company in San Francisco seeks a Research Scientist to develop and enhance models for audiovisual data processing. This role involves researching novel methods for audio and video enhancement, running large-scale experiments, and collaborating with engineers for workflow integration. Ideal candidates have a strong background in deep learning, proficiency in PyTorch, and familiarity with signal processing. The environment is fast-paced, rewarding impactful applied research efforts.

Qualifications

  • Experience training models for audio or video processing.
  • Solid understanding of signal processing fundamentals.
  • Ability to run experiments at scale.

Responsibilities

  • Research and develop models for audio and video enhancement.
  • Experiment with architectures and data augmentation.
  • Build evaluation frameworks to measure model performance.
  • Collaborate with engineers for model integration.

Skills

Deep learning research background
Proficiency in PyTorch
Experience in signal processing
Publication track record
Ability to adapt in early-stage environments

Job description

Velvet is a data research company building the datasets that power the next generation of multimodal AI. Founded by Lucas Mantovani (ex Meta FAIR) and Lucas Tucker (ex Adobe Infra), our mission is to make AI more human by producing high-quality audiovisual training data for frontier labs.

We’re hiring a Research Scientist to develop and fine‑tune models for video and audio data processing and enhancement, as well as to conduct data‑oriented research that pushes the boundaries of multimodal quality.

What You’ll Do
  • Research, develop, and fine‑tune models for audio and video enhancement — including denoising, super‑resolution, speech restoration, and perceptual quality improvement — ensuring outputs meet the standards required for frontier model training.
  • Experiment with novel architectures, training objectives, and data augmentation strategies to improve model performance across diverse and noisy real‑world audiovisual data.
  • Build evaluation frameworks and benchmarks to rigorously measure enhancement quality, guiding iterative model improvement.
  • Collaborate with infrastructure and data pipeline engineers to integrate trained models into large‑scale processing workflows that handle wide variation in speech, visual quality, and format.
What We’re Looking For
  • Strong research background in deep learning, with hands‑on experience training and fine‑tuning models for audio processing, video processing, or related domains.
  • Proficiency in PyTorch and experience designing and running experiments at scale.
  • Solid understanding of signal processing fundamentals and how they inform model design for enhancement tasks.
  • A publication track record or demonstrated research output in relevant areas (audio/speech enhancement, video restoration, generative models, multimodal learning).
  • Ability to work effectively in an early‑stage environment where scope is broad and priorities shift fast.
Even Better
  • Prior work at a frontier AI lab or data company focused on multimodal data.
  • Experience fine‑tuning large‑pretrained models (diffusion models, autoencoders, or transformer‑based architectures) for perceptual quality tasks.
  • Familiarity with perceptual quality metrics and human evaluation methodologies for audio and video.
  • Track record working with datasets spanning tens of thousands of hours of audio or video.
You’ll Thrive Here If
  • You’re excited by applied research with immediate, visible impact on data quality and downstream model performance.
  • You move fluidly between reading papers, writing training loops, and analyzing failure cases.
  • You hold yourself to a high bar for rigor — because you understand that model quality directly determines the value of the data we produce.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Machine Learning Engineer
Founding Machine Learning Engineer

Velvet • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Engineer, Data
Research Engineer, Data

Harnham • California (MO)

On-site
USD 100,000 - 130,000
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

synthesia • United States

On-site
USD 180,000 - 260,000
Research Engineer
Research Engineer

Goliath Partners Inc. • New York (NY)

On-site
USD 120,000 - 180,000
Multimodal Audio-Video Enhancement Scientist
Multimodal Audio-Video Enhancement Scientist

Velvet • San Francisco (CA)

On-site
USD 110,000 - 150,000
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research, Audio Expertise
Research, Audio Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Scientist, Video Understanding & World Models
Research Scientist, Video Understanding & World Models

Kindredventures • New York (NY)

On-site
USD 120,000 - 150,000
Research Scientist
Research Scientist

David AI • San Francisco (CA)

On-site
USD 180,000 - 250,000
Unlimited PTO
Health, dental, and vision coverage
FSA & HSA access
+3
Research, Audio Expertise
Research, Audio Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1