Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Nuance Labs in Seattle is seeking an experienced human-evaluation scientist to turn subjective judgments into measurable signals for real-time AI avatars. You will design studies, define rubrics, and coordinate with researchers to align outcomes with product goals.
You will own the end-to-end evaluation pipeline, from study design to data analysis, and you’ll collaborate directly with founders and the modeling team to decide which models ship next.
Nuance Labs is building photorealistic, real-time AI avatars with emotional intelligence: a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
We're a research company, with PhDs from MIT, UW, Oxford, CMU, and Johns Hopkins, and industry experience from Apple, Meta, Amazon AGI, and more. Backed by Accel, Lightspeed, South Park Commons, and NVIDIA, we combine frontier research with ruthless engineering needed for consumer-grade, real-time systems. The team is small, the work is real, and the problems are unsolved.
Most conversational AI avatars today are hacks — a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time. Current systems take 2-5 seconds to respond; natural conversation requires sub-500ms. That's a 10x improvement, and it demands rethinking the entire stack.
That rethinking starts with full-duplex: an AI that listens and speaks simultaneously, perceives emotion in real time, and responds with a face that actually reflects it. It's an extremely hard problem, and we're developing foundation models designed for it from the ground up.
"Does this avatar feel human?" is the question our whole company is organized around — and no automated metric can answer it. Lip-sync error and video quality scores say nothing about whether a smile landed as sincere or unsettling, whether a conversation felt warm or hollow, or whether someone would want to talk to our avatar again tomorrow.
Your job is to turn human judgment into a reliable signal our researchers can train and ship against. You'll take the most ambiguous problems in our field (naturalness, emotional resonance, trust, presence) and design studies whose results people actually agree on. When two models differ, your study is the tiebreaker. When a model "feels off" and nobody can say why, your investigation finds the cause.
This is a hands-on IC role and our first hire dedicated to human evaluation. You'll own it end to end (what to measure, how to measure it, who rates it, and what the results mean) working directly with the founders and the modeling team. Your findings will decide which models ship and what we train next.
No candidate checks every box. If the "good fit" list sounds like you but your background is unconventional, we'd like to hear from you anyway.
$160,000 - $190,000 base salary, plus meaningful equity. We think long-term ownership matters and structure equity accordingly.