Evaluations Lead, Generative Speech AI

Sanas

Palo Alto (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Sanas is building a full Speech AI suite, all working together as one platform. As Evaluations Lead, you’ll design evaluation frameworks and benchmarking systems that answer how models are evaluated at scale, sitting at the intersection of research, product, and infrastructure to build metrics, systems, and studies across all products.

This role blends scientific rigor with hands-on execution, ensuring progress is measured by real-world qualities like understanding, naturalness, and adaptability

Qualifications

  • Experience designing or implementing evaluation frameworks for generative models — audio, text, or multimodal.
  • Ability to turn open-ended research ideas into production-ready systems.
  • Creative skills in defining quantitative metrics for inherently subjective qualities.

Responsibilities

  • Identify and define the model capabilities and behaviors that matter for evaluation.
  • Build and ship evaluation pipelines with robust statistical analysis and clear reporting.
  • Partner directly with model training and research teams to embed evaluation into the development loop.
  • Prototype new user studies and behavioral experiments that ground evaluation in real-world use.

Skills

Evaluation frameworks
Generative models
Statistical analysis
Experiment design
Production-ready system

Job description

Sanas is building a full Speech AI suite, all working together as one platform. As Evaluations Lead, you’ll design evaluation frameworks and benchmarking systems that answer how models are evaluated at scale, sitting at the intersection of research, product, and infrastructure to build metrics, systems, and studies across all products.

This role blends scientific rigor with hands-on execution, ensuring progress is measured by real-world qualities like understanding, naturalness, and adaptability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Research Evaluations
Member of Technical Staff, Research Evaluations

Sanas • Palo Alto (CA)

On-site
USD 140,000 - 190,000
Research Scientist (Model Evaluation)
Research Scientist (Model Evaluation)

Jobless • Palo Alto (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Research Scientist (Model Evaluation)
Research Scientist (Model Evaluation)

Sanas.AI Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 140,000 - 230,000
Research Scientist (Model Evaluation)
Research Scientist (Model Evaluation)

Sanas • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Research Scientist: Model Evaluation & Quality Metrics
Research Scientist: Model Evaluation & Quality Metrics

Sanas • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Research Scientist: Model Evaluation & Metrics
Research Scientist: Model Evaluation & Metrics

Sanas.AI Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 140,000 - 230,000
Head of AI Evaluation & Benchmarks
Head of AI Evaluation & Benchmarks

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 225,000 - 275,000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
Senior AI Systems Engineer: Model Evaluation & QA
Senior AI Systems Engineer: Model Evaluation & QA

Deepgram • United States

Remote
USD 140,000 - 210,000
Lead Specialist, AI Scientist
Lead Specialist, AI Scientist

Pearson • Town of Poland (NY)

On-site
USD 90,000 - 130,000
Member of Technical Staff, ML Inference Engineering
Member of Technical Staff, ML Inference Engineering

Sanas • Palo Alto (CA)

On-site
USD 180,000 - 240,000