AI Research Scientist

Thespian Labs

Somerville, Northern (MA, KY)

Hybrid

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Equal opportunity employer

Job summary

Thespian Labs is seeking an AI Research Scientist to design and implement state-of-the-art generative models that power embodied behavior. You will work with diffusion-based and transformer architectures to synthesize 3D content, motion, and audio for real-time control of humanoids and digital characters.

You will own research problems end to end, collaborating with data and engineering teams to turn large-scale datasets into observable behavior.

Qualifications

  • Proven experience developing and training deep learning models, particularly transformers and/or diffusion models.
  • Strong programming skills in Python and proficiency with modern ML frameworks like PyTorch.
  • Solid understanding of 3D computer vision and geometry, including how shape, motion, and space are represented and generated.
  • A strong portfolio of publications at top-tier conferences such as CVPR, ICCV, or NeurIPS.
  • Ability to take a model from research to a system that runs, with care for quality, performance, and scale.

Responsibilities

  • Develop and train advanced generative models, particularly diffusion-based architectures, for high-fidelity 3D content in real time.
  • Explore and implement cutting-edge techniques across generative modeling, including visual, motion, and audio components.
  • Design the system to combine multiple model types into a single forward pass rather than a pipeline.
  • Collaborate with data engineering to define data requirements and use large-scale datasets for training.
  • Optimize models for quality, performance, and scalability to run in live scenarios and on robots.
  • Stay current with the latest research in generative AI and 3D computer vision and bring findings into work quickly.
  • Take research from idea to deployment, owning problems end to end with real humanoids and live digital characters.

Skills

Deep learning
Transformers
Diffusion models
Python
3D computer vision
Research publications
End-to-end deployment

Tools

PyTorch

Job description

Research the generative models at the core of one foundation model for embodied behavior. One model, any body, from a single forward pass.

Thespian Labs is an embodied intelligence research lab with roots at MIT and deep expertise in AI. We are building a foundation model for behavior, the layer between reasoning and execution where a body has to do something coherent while the world keeps moving. Reasoning can plan and decide. Execution can render pixels and drive motors. The layer in between is the one nobody has built, and it is the whole of our work.

We believe intelligence is something a body does in the world. A screen can answer you, but a body can be with you. That is why we are building one model that drives any body from a single forward pass, with perception, action, expression, and voice in a single stream. If language was the last decade of this work, embodiment is the next.

About the Role

As an AI Research Scientist at Thespian Labs, you will research, design, and implement state-of-the-art generative models that sit at the core of our platform, the models that synthesize an embodied performance from a single forward pass. Working with novel, often diffusion-based and transformer architectures, you will generate the complex, high-fidelity 3D content, along with the motion, visual, and audio components, that let one model drive a humanoid on a real floor or a digital character on a screen.

This is core technology, not a wrapper around someone else’s. The team is small and the loop is short. You will own research problems end to end, from architecture to a model running on a real body, and work closely with the data team to turn unique, large-scale datasets into behavior the world can see. We move fast, we stay close to the latest research, and we hold the work to a high bar. Not a benchmark, but the world.

Responsibilities
  • Develop and train advanced generative models, particularly diffusion-based architectures, for the synthesis of dynamic, high-fidelity 3D content that drives a body in real time.
  • Explore and implement cutting-edge techniques across generative modeling, including the visual, motion, and audio generation models that come together as a single embodied performance.
  • Design the system so these components work as one, composing multiple model types into a single forward pass rather than a pipeline of disconnected parts.
  • Collaborate closely with the data engineering team to define data requirements and leverage unique, large-scale datasets for model training.
  • Optimize models for quality, performance, and scalability, so research-grade results run fast enough to act in a live room and on a real robot.
  • Stay current with the latest research in generative AI and 3D computer vision, and bring new findings into our work quickly.
  • Take research from idea to deployment, owning problems end to end and closing the loop onto real humanoids and live digital characters.
What We’re Looking For
  • Proven experience developing and training deep learning models, particularly transformers and/or diffusion models.
  • Strong programming skills in Python and proficiency with modern ML frameworks like PyTorch.
  • A solid understanding of 3D computer vision and geometry, including how shape, motion, and space are represented and generated.
  • A strong portfolio of relevant projects or publications at top-tier conferences such as CVPR, ICCV, or NeurIPS.
  • The ability to take a model from research to a system that runs, with care for quality, performance, and scale rather than just a result in a notebook.
  • A bias for ownership and the appetite to drive open, ambiguous problems end to end in a small team.
Preferred Qualifications
  • Familiarity with audio, motion, or other generative domains beyond your core specialization.
  • Experience working on multiple model types within a single system, and making them cohere.
  • Exposure to audio or speech generation models, which are valuable here though not a primary specialization.
  • Hands-on experience with 3D data formats and representations such as meshes, point clouds, and implicit fields.
  • A research background with publications in venues like CVPR, ICCV, or NeurIPS.
  • Experience with the realities of real-time or interactive generation, where model output has to keep up with a live scene.

Tell us how you think: papers, code, demos, or a project you couldn't put down. We read everything.

Thespian Labs is an equal-opportunity employer. We hire for the work, and we want a lab full of people who don't all think alike. We sponsor visas for the right person.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Robotics Research Scientist
Robotics Research Scientist

Thespian Labs • Somerville (MA), Northern (KY)

Hybrid
USD 120,000 - 190,000
Visa sponsorship
Equal opportunity employer
VP of Engineering
VP of Engineering

Thespian Labs • Somerville (MA), Northern (KY)

Hybrid
USD 220,000 - 300,000
Visa sponsorship
Embodied AI Research Scientist — Real-Time 3D Synthesis
Embodied AI Research Scientist — Real-Time 3D Synthesis

Thespian Labs • Somerville (MA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Visa sponsorship
Equal opportunity employer
Member of Technical Staff, Research
Member of Technical Staff, Research

Odyssey • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Research Scientist - Interactive Avatars
Research Scientist - Interactive Avatars

Synthesia • United States

Remote
USD 170,000 - 230,000
Member of Technical Staff, Applied Research
Member of Technical Staff, Applied Research

Odyssey • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

synthesia • United States

On-site
USD 180,000 - 260,000
Research Scientist - Video Diffusion
Research Scientist - Video Diffusion

Nuance Labs • Seattle (WA)

On-site
USD 150,000 - 210,000
AI Research Scientist
AI Research Scientist

Noösphere • Seattle (WA)

On-site
USD 150,000 - 210,000
Early-stage equity
Benefits package
MOTS - World Models
MOTS - World Models

Alexander Chapman Ltd • Palo Alto (CA)

On-site
USD 150,000 - 190,000