This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Research Scientist – Interactive Avatars based in Germany.
This is a senior research role at the forefront of generative AI, focused on building the next generation of interactive, avatar-based video agents. You’ll work within a multidisciplinary R&D organization combining research, engineering, and data expertise. Your focus will be on creating models that can understand audio and visual signals and respond with natural, human-like conversational behavior. You’ll shape research direction while balancing ambitious long-term opportunities with near-term product impact. The role combines deep technical research, hands-on model development, evaluation, and research leadership. You’ll also mentor researchers and influence technical decisions across multiple teams, helping transform breakthrough ideas into production-ready capabilities.
Accountabilities
- Define the research direction and roadmap for dyadic interaction modeling, balancing long-term research opportunities with immediate product priorities.
- Advance the perceptual capabilities of interactive AI agents, including understanding user audio and video and generating contextually appropriate responses.
- Develop and post-train multimodal models capable of producing rich, natural, and responsive dyadic interactions from audio and video inputs.
- Adapt and extend diffusion models to incorporate conversational conditioning signals such as dialogue state, turn‑taking, listener cues, and contextual reactions.
- Research and develop models capable of generating natural conversational behaviors, including gaze, facial expressions, body reactions, and backchannel responses.
- Design robust evaluation frameworks, benchmarks, and automated test suites to continuously measure interaction quality.
- Collaborate closely with data specialists to define data requirements and shape high‑quality datasets for model development and training.
- Lead technical decisions spanning research, data, and engineering, ensuring strong alignment between research goals and product requirements.
- Lead and mentor a small team of researchers, supporting their technical development and helping establish a high‑performing research culture.
- Drive research projects from initial hypothesis and experimentation through validation, productionization, and product impact.
- Communicate research hypotheses, methodologies, experimental results, and technical recommendations clearly to both research and cross‑functional audiences.
Requirements
- Deep expertise in machine learning with extensive hands‑on experience working with diffusion models, ideally applied to video generation, avatar generation, or related multimodal applications.
- Strong research track record demonstrated through publications at leading conferences such as CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or SIGGRAPH, or equivalent evidence of significant research impact.
- Expertise in areas such as video diffusion, world models, dyadic interaction modeling, multimodal generation, or related fields.
- Proven experience leading a small team of researchers and mentoring junior research talent.
- Demonstrated ability to take research concepts from initial idea through experimentation and into production systems.
- Strong proficiency in PyTorch and modern machine learning tooling for large‑scale model training.
- Strong understanding of multimodal learning, generative modeling, and large‑scale ML experimentation.
- Excellent analytical and problem‑solving abilities, with a rigorous approach to designing experiments and evaluating results.
- Clear communication skills and the ability to explain complex technical concepts, influence research direction, and collaborate effectively across teams.
- Strong ownership and ability to operate in a fast‑moving research environment where experimentation and iteration are encouraged.
Nice to have
- Experience with real‑time or streaming generation, including autoregressive video diffusion.
- Expertise in model distillation or other techniques designed to achieve low‑latency inference.
- Experience with audio‑driven facial animation, gesture generation, or full‑body motion modeling.
- Experience in conversational modeling, including turn‑taking, backchanneling, listener‑response generation, or related interaction behaviors.
- Experience developing AI systems designed for real‑time human‑machine interaction.
Benefits
- Competitive compensation package.
- Fully remote working option within Europe.
- Hybrid working opportunities for employees based near London, Munich, or Zurich offices.
- 25 days of annual leave plus public holidays.
- Opportunity to work alongside a large team of AI researchers and engineers at the forefront of generative AI.
- Direct opportunity to influence the research roadmap and development of emerging interactive AI technologies.
- Strong culture focused on building, experimentation, autonomy, and high‑impact execution.
- Regular opportunities to connect with colleagues through office hubs, planning sessions, and social events.
- Collaborative environment with access to multidisciplinary expertise across research, engineering, and data.
- Opportunity to mentor researchers and contribute to the development of advanced AI research capabilities.
- Work on real‑world AI products used by leading organizations across multiple industries.
- Commitment to responsible AI, with strong emphasis on safety, ethics, security, and human‑centered technology.
- Additional location‑specific benefits depending on where you are based.