Qualitative Evaluation Engineer

Luma

United States

On-site

USD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma seeks a Qualitative Evaluation Engineer to define believability, identity retention, and scene coherence, building evaluative systems that match human perception. You will work with researchers and technical artists to steer model development beyond checkbox metrics.

In the first 90 days you will immerse, stand up a qualitative framework, and scale it into the eval loop, turning nuanced judgments into actionable feedback for fine-tuning and dataset curation.

Qualifications

  • 5+ years in product evaluation, UX research, model testing, or similar qualitative assessment.
  • Master's or higher in Cognitive Science, HCI, Design Research, Psychology, Media Studies, or related field.
  • Deep familiarity with creative workflows for generative models (animation, filmmaking, digital art, VFX).
  • Systems thinking: you can define abstract qualities like believability or scene coherence in clear evaluative terms.
  • Excellent written communication and the ability to synthesize nuanced judgment into actionable insight.
  • Comfort working across engineers, researchers, and creatives.

Responsibilities

  • Evaluate generative model performance across diverse tasks, prompts, and modalities, and surface the failure modes, regressions, and edge cases that hurt product quality.
  • Build and maintain qualitative evaluation frameworks that are scalable and reusable.
  • Translate high-level product goals into concrete evaluative criteria.
  • Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluations.
  • Turn nuanced judgments into clear feedback that informs fine-tuning, dataset curation, and product UX.
  • Work closely with technical artists and engineers to keep evaluations aligned with model capabilities and real use cases.

Skills

Qualitative evaluation
UX research
Model testing
Written communication
Cross-functional work

Education

Master's degree or higher in Cognitive Science / HCI / Design Research / Psychology / Media Studies

Job description

You'll own how Luma judges whether its models are actually good, past the point where numbers stop telling the story. As our Qualitative Evaluation Engineer, you'll build the frameworks that pin down believability, identity retention, and scene coherence, and turn them into insight that steers model development. This isn't a checkbox-metrics role. You're building evaluative systems that match the messiness of human perception and creative intent, and much of that framework doesn't exist yet. It fits someone who can take a fuzzy quality and define it in clear, testable terms, working shoulder to shoulder with researchers and technical artists. If you want work scored purely on quantitative dashboards, this isn't it.

What You'll Own
  • Evaluate generative model performance across diverse tasks, prompts, and modalities, and surface the failure modes, regressions, and edge cases that hurt product quality.
  • Build and maintain qualitative evaluation frameworks that are scalable and reusable.
  • Translate high-level product goals into concrete evaluative criteria.
  • Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluations.
  • Turn nuanced judgments into clear feedback that informs fine-tuning, dataset curation, and product UX.
  • Work closely with technical artists and engineers to keep evaluations aligned with model capabilities and real use cases.
First 90 Days

One way the first 90 could unfold.

  • Days 1–30 — Immerse & Diagnose: Learn the models and the creative use cases they serve, and audit how quality is judged today and where it misses.
  • Days 30–60 — Ship & Validate: Stand up a qualitative framework for one high-priority capability and run it on real outputs, producing insight the team acts on.
  • Days 60–90 — Scale & Systemize: Make the framework reusable across capabilities and wire it into the model/data/eval loop.
What You Bring
  • 5+ years in product evaluation, UX research, model testing, or similar structured qualitative assessment.
  • Master's or higher in Cognitive Science, HCI, Design Research, Psychology, Media Studies, or a related field.
  • Deep familiarity with creative workflows for generative models (animation, filmmaking, digital art, VFX).
  • Systems thinking: you can define abstract qualities like believability or scene coherence in clear evaluative terms.
  • Excellent written communication and the ability to synthesize nuanced judgment into actionable insight.
  • Comfort working across engineers, researchers, and creatives.
Nice to Have
  • Background in motion, visual effects, or storytelling pipelines.
  • Experience evaluating AI-generated media (video, images, 3D).
  • Prior work building internal tools for qualitative data collection or scoring.
  • Familiarity with prompt engineering and reference-based inputs.

About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision.

Luma is an equal opportunity employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Qualitative Evaluation Engineer, Multimodal AI
Qualitative Evaluation Engineer, Multimodal AI

Luma AI • United States

Remote
USD 120,000 - 170,000
Research Engineer - Evaluations
Research Engineer - Evaluations

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Qualitative Evaluation Engineer: Define Believability
Qualitative Evaluation Engineer: Define Believability

Luma • United States

On-site
USD 120,000 - 190,000
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000
Senior Multimodal AI Evaluation Engineer
Senior Multimodal AI Evaluation Engineer

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Senior LLM Evaluation Engineer
Senior LLM Evaluation Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
Staff Product Software Engineer
Staff Product Software Engineer

Luma • Redwood City (CA)

On-site
USD 260,000 - 380,000
Product Manager, Applied Research
Product Manager, Applied Research

Luma • Redwood City (CA)

On-site
USD 225,000 - 325,000
Visual Designer, Product
Visual Designer, Product

Ch • Redwood City (CA), Northern (KY)

Hybrid
USD 110,000 - 170,000
Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E
Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E

Apple • San Diego (CA)

On-site
USD 140,000 - 190,000