An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Descript's Research team builds models behind the product's standout features, including Video Regenerate and lipsync. This role focuses on multimodal understanding, training models to perceive edited media the way a human editor does, enabling better evaluation and smarter editing decisions.
The candidate will design and implement deep learning solutions, evaluate experiments, and help ship production features.
Descript's Research team builds the models behind the product's most distinctive features: Video Regenerate and lipsync, video translation, zero-shot voice and roomtone cloning, and Studio Sound. We don't build general-purpose generative models. We pick specific problems in the editing workflow and build specialized models for them. This isn't research for its own sake. Everything we build is meant to ship, and most of it has, going from prototype to a production feature used by millions of creators within months.
This role is focused on multimodal understanding: training models to perceive edited media the way a human video editor does. Underlord, our AI editing agent, reasons about a project largely through a textual representation of it. Giving it direct perception of the media it's working on is what will let it judge its own output and reason about the creative choices in an edit, not just the structure of a project. It's also an open research problem, since there's no settled way to represent or evaluate editorial craft, whether a cut lands or whether the pacing works. We have a unique dataset to work with.
At least one of the following must be true:
More senior candidates (Senior and Staff) should also bring a track record of owning research direction rather than executing a plan handed to them, and experience mentoring or technically leading other researchers or engineers.
Direct experience in multimodal understanding is welcome but not required, and we don't require domain-specific expertise in computer vision or speech and audio. Our team spans both, and strong general deep learning ability transfers. We hire against the bar above, and then expect you to grow into the domain. Depth in any of these is a strong signal:
Base salary range: $197,000-$262,500, plus equity and benefits. Final offer amounts will carefully consider multiple factors, including prior experience, expertise, location, and level, and may vary from the amount above.
Descript is an equal opportunity workplace-we are dedicated to equal employment opportunities regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, or Veteran status. We believe in actively building a team rich in diverse backgrounds, experiences, and opinions to better allow our employees, products, and community to thrive.