Qualitative Evaluation Engineer: Define Believability

Luma

United States

On-site

USD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma seeks a Qualitative Evaluation Engineer to define believability, identity retention, and scene coherence, building evaluative systems that match human perception. You will work with researchers and technical artists to steer model development beyond checkbox metrics.

In the first 90 days you will immerse, stand up a qualitative framework, and scale it into the eval loop, turning nuanced judgments into actionable feedback for fine-tuning and dataset curation.

Qualifications

  • 5+ years in product evaluation, UX research, model testing, or similar qualitative assessment.
  • Master's or higher in Cognitive Science, HCI, Design Research, Psychology, Media Studies, or related field.
  • Deep familiarity with creative workflows for generative models (animation, filmmaking, digital art, VFX).
  • Systems thinking: you can define abstract qualities like believability or scene coherence in clear evaluative terms.
  • Excellent written communication and the ability to synthesize nuanced judgment into actionable insight.
  • Comfort working across engineers, researchers, and creatives.

Responsibilities

  • Evaluate generative model performance across diverse tasks, prompts, and modalities, and surface the failure modes, regressions, and edge cases that hurt product quality.
  • Build and maintain qualitative evaluation frameworks that are scalable and reusable.
  • Translate high-level product goals into concrete evaluative criteria.
  • Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluations.
  • Turn nuanced judgments into clear feedback that informs fine-tuning, dataset curation, and product UX.
  • Work closely with technical artists and engineers to keep evaluations aligned with model capabilities and real use cases.

Skills

Qualitative evaluation
UX research
Model testing
Written communication
Cross-functional work

Education

Master's degree or higher in Cognitive Science / HCI / Design Research / Psychology / Media Studies

Job description

Luma seeks a Qualitative Evaluation Engineer to define believability, identity retention, and scene coherence, building evaluative systems that match human perception. You will work with researchers and technical artists to steer model development beyond checkbox metrics.

In the first 90 days you will immerse, stand up a qualitative framework, and scale it into the eval loop, turning nuanced judgments into actionable feedback for fine-tuning and dataset curation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Qualitative Evaluation Engineer, Multimodal AI
Qualitative Evaluation Engineer, Multimodal AI

Luma AI • United States

Remote
USD 120,000 - 170,000
Qualitative Evaluation Engineer
Qualitative Evaluation Engineer

Luma • United States

On-site
USD 120,000 - 190,000
Research Engineer - Evaluations
Research Engineer - Evaluations

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Senior Multimodal AI Evaluation Engineer
Senior Multimodal AI Evaluation Engineer

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000
AI Evaluation QA Engineer: Scale Evals & Improve Answers
AI Evaluation QA Engineer: Scale Evals & Improve Answers

Socket.dev • New York (NY), Austin (TX), Miami (FL)

On-site
USD 110,000 - 160,000
AI Evaluation Engineer: Scale QA for LLMs
AI Evaluation Engineer: Scale QA for LLMs

Appnovation • Dallas (TX)

On-site
USD 95,000 - 140,000
Eval Infrastructure Engineer — Define LLM Quality Metrics
Eval Infrastructure Engineer — Define LLM Quality Metrics

Firecrawl • San Francisco (CA)

Hybrid
USD 160,000 - 240,000
Unlimited PTO
12 weeks fully paid parental leave
Wellness stipend
+5
AI Evaluation Scientist: Trust & Safety Leader
AI Evaluation Scientist: Trust & Safety Leader

Steampunk • McLean (VA)

Hybrid
USD 105,000 - 145,000
Evals Lead
Evals Lead

Fluency Digital, Inc. • New York (NY)

On-site
USD 120,000 - 150,000