An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Cantina Labs is hiring a Research / ML Engineer to join our Speech Team and own the audio side of multimodal generation. You will design, train, and deploy audio VAEs, neural codecs, and diffusion-based backbones, with multi-speaker conditioning and AV synchronization for cinematic dialogue and sound design.
You will collaborate across research, video, data, and infra to ship reliable, cost-aware models; lead experiments, tune data pipelines, and push safe, steerable AI that speaks and emotes in
Cantina Labs is hiring a Research / ML Engineer to join our Speech Team and own the audio side of multimodal generation. You will design, train, and deploy audio VAEs, neural codecs, and diffusion-based backbones, with multi-speaker conditioning and AV synchronization for cinematic dialogue and sound design.
You will collaborate across research, video, data, and infra to ship reliable, cost-aware models; lead experiments, tune data pipelines, and push safe, steerable AI that speaks and emotes in