Applied Research Scientist, AI Research

Descript

San Francisco (CA)

On-site

USD 197,000 - 263,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

Descript's Research team builds models behind the product's standout features, including Video Regenerate and lipsync. This role focuses on multimodal understanding, training models to perceive edited media the way a human editor does, enabling better evaluation and smarter editing decisions.

The candidate will design and implement deep learning solutions, evaluate experiments, and help ship production features.

Qualifications

  • Proven ability to design and implement deep learning algorithms.
  • Strong programming skills and fluency in PyTorch.
  • Track record of generating new ideas and shipping results.

Responsibilities

  • Train models from scratch or fine-tune existing foundation models.
  • Design benchmarks and evaluation methods for edited media.
  • Ship production features in collaboration with engineering teams.
  • Mentor or lead researchers as needed for senior roles.

Skills

Deep learning
PyTorch
Experimental judgment
Communication
Publications

Education

PhD or Master's in deep learning or related field

Job description

Descript's Research team builds the models behind the product's most distinctive features: Video Regenerate and lipsync, video translation, zero-shot voice and roomtone cloning, and Studio Sound. We don't build general-purpose generative models. We pick specific problems in the editing workflow and build specialized models for them. This isn't research for its own sake. Everything we build is meant to ship, and most of it has, going from prototype to a production feature used by millions of creators within months.

This role is focused on multimodal understanding: training models to perceive edited media the way a human video editor does. Underlord, our AI editing agent, reasons about a project largely through a textual representation of it. Giving it direct perception of the media it's working on is what will let it judge its own output and reason about the creative choices in an edit, not just the structure of a project. It's also an open research problem, since there's no settled way to represent or evaluate editorial craft, whether a cut lands or whether the pacing works. We have a unique dataset to work with.

  • Video Regenerate : regenerating a speaker's lower face to match new or translated audio
  • Jumpcut Smoothing : generating a bridge across a cut so the join plays like a continuous take
  • PoDAR : disentangling power from semantics in audio latents to make them easier to model
  • Multimodal understanding: build vision-language systems that let Descript's agentic editing features reason over the visual and audio content of a project.
  • Evaluation: design the benchmarks and evals that make editorial quality measurable, and that balance quality against cost and latency.
  • Data: build the datasets your work depends on, including synthetic data generation where real examples don't exist at scale.
  • Training: train specialized models from scratch or fine-tune existing foundation models, whichever gets the capability we need.
  • Shipping: take models from prototype to production with the agent and engineering teams.
  • Direction-setting: identify the next research direction that should become a Descript feature, not just a paper. More senior candidates should expect to own this directly; more junior candidates will grow into it.
  • Publishing: take your work to academic venues if you'd like. We support it, but it isn't a requirement of the role.
What you bring
Required
  • Proven ability to design and implement deep learning algorithms, demonstrated by publications, open-source work, or models you've shipped.
  • Strong programming skills and deep fluency in PyTorch.
  • A track record of generating new ideas in machine learning. You produce more ideas than you can implement, and once an experiment setup is established, you can run and evaluate many of them quickly rather than being bottlenecked on infrastructure.
  • Strong experimental judgment. You test ideas fast, and you're honest with yourself and the team about which ones don't pan out.
  • Clear written and verbal communication, including when a direction isn't working, so the team doesn't waste time following a lead that's already dead.
  • A PhD or Master's in deep learning or a related field, or equivalent experience. We care about the track record more than the credential.

At least one of the following must be true:

  • Lead or first author of an accepted publication in a top venue: CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or similar.
  • Played a key role in shipping a production feature with deep learning as a core component.

More senior candidates (Senior and Staff) should also bring a track record of owning research direction rather than executing a plan handed to them, and experience mentoring or technically leading other researchers or engineers.

Where breadth helps

Direct experience in multimodal understanding is welcome but not required, and we don't require domain-specific expertise in computer vision or speech and audio. Our team spans both, and strong general deep learning ability transfers. We hire against the bar above, and then expect you to grow into the domain. Depth in any of these is a strong signal:

  • Vision-language models and multimodal understanding.
  • Generative modeling for video, audio, or images.
  • Post-training, fine-tuning, and RL on large foundation models.
  • Building evaluation systems for generative or agentic outputs where metrics resist clean definitions.
  • Taking a research idea through to a shipped, production-facing feature.
Compensation and benefits

Base salary range: $197,000-$262,500, plus equity and benefits. Final offer amounts will carefully consider multiple factors, including prior experience, expertise, location, and level, and may vary from the amount above.

Descript is an equal opportunity workplace-we are dedicated to equal employment opportunities regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, or Veteran status. We believe in actively building a team rich in diverse backgrounds, experiences, and opinions to better allow our employees, products, and community to thrive.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied Research Scientist, AI Research Descript San Francisco, CA or Remote, US $197,000 - $262,500/yr
Applied Research Scientist, AI Research Descript San Francisco, CA or Remote, US $197,000 - $262,500/yr

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 197,000 - 263,000
Healthcare package
401k matching
Catered lunches
+1
Software Engineer, Infrastructure Descript San Francisco, CA or Remote, US $220,000 - $292,000/yr
Software Engineer, Infrastructure Descript San Francisco, CA or Remote, US $220,000 - $292,000/yr

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 292,000
Software Engineer, Product
Software Engineer, Product

descript • United States

On-site
USD 220,000 - 265,000
Senior Software Engineer, Product
Senior Software Engineer, Product

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 265,000
Generous healthcare package
401k matching
Catered lunches
+1
Software Engineer, Product
Software Engineer, Product

Descript • San Francisco (CA)

Hybrid
USD 220,000 - 265,000
Healthcare package
401(k) matching
Catered lunches
+1
Software Engineer, Product Descript San Francisco, CA or Remote, US $220,000 - $265,000/yr
Software Engineer, Product Descript San Francisco, CA or Remote, US $220,000 - $265,000/yr

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 265,000
Healthcare package
401k matching
Catered lunches
+1
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Descript • San Francisco (CA)

Hybrid
USD 220,000 - 292,000
Catered lunches
Flexible vacation time
401k matching
+1
Engineering Manager (Agent)
Engineering Manager (Agent)

Descript • San Francisco (CA)

On-site
USD 222,431 - 261,684
Generous healthcare package
401k matching program
Catered lunches
+1
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

synthesia • United States

On-site
USD 180,000 - 260,000
Research Staff, Data Science
Research Staff, Data Science

Madrona Venture Labs • United States

On-site
USD 150,000 - 220,000
Equity offerings
Annual bonuses