Senior Multimodal AI Evaluation Engineer

Luma AI

San Francisco, New York (CA, NY)

On-site

USD 170,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma AI is seeking a Research Engineer to design and scale the infrastructure powering model evaluation efforts. You will build pipelines, metrics, and automated systems that connect model output, evaluation, and improvement across research, engineering, and product teams.

Ideal candidates have 5+ years in ML evaluation, strong Python skills, and experience with visual data and multimodal models. The role emphasizes scalable systems, CI/CD, and measurement-driven development.

Qualifications

  • Master's or PhD in Computer Science, Machine Learning, or a related technical field (or equivalent industry experience).
  • 5+ years of experience building ML evaluation systems, model pipelines, or large-scale infrastructure.
  • Hands-on experience working with visual data (images and/or video), including evaluation, modeling, or data preparation.
  • Proficiency in Python and ML frameworks (PyTorch, JAX, or TensorFlow).
  • Familiarity with human-in-the-loop evaluation workflows and how to scale them with automation.
  • Strong background in machine learning, with experience in generative models (diffusion, LLMs, multimodal architectures).
  • Strong software engineering skills (CI/CD, testing, data pipelines, distributed systems).

Responsibilities

  • Design and implement scalable pipelines for automated evaluation of generative models, with a focus on visual and multimodal outputs (image, video, text, audio).
  • Develop novel metrics and evaluation models that capture qualities like fidelity, coherence, temporal consistency, and alignment with human intent.
  • Integrate evaluation signals into training loops (including reinforcement learning and reward modeling) to continuously improve model performance.
  • Build infrastructure for large-scale regression testing, benchmarking, and monitoring of multimodal generative models.
  • Collaborate with researchers running human studies to translate human evaluation frameworks into automated or semi-automated systems.
  • Partner with model researchers to identify failure cases and build targeted evaluation harnesses.
  • Maintain dashboards, reporting tools, and alerting systems to surface evaluation results to stakeholders.
  • Stay current with emerging evaluation techniques in generative AI, multimodal LLMs, and perceptual quality assessment.

Skills

Python
CI/CD
Distributed systems
ML frameworks (PyTorch/JAX/TF)

Education

Master's or PhD in Computer Science, ML, or related field

Tools

PyTorch
JAX
TensorFlow

Job description

Luma AI is seeking a Research Engineer to design and scale the infrastructure powering model evaluation efforts. You will build pipelines, metrics, and automated systems that connect model output, evaluation, and improvement across research, engineering, and product teams.

Ideal candidates have 5+ years in ML evaluation, strong Python skills, and experience with visual data and multimodal models. The role emphasizes scalable systems, CI/CD, and measurement-driven development.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer - Evaluations
Research Engineer - Evaluations

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Inference Systems Engineer for Scalable Multimodal AI
Inference Systems Engineer for Scalable Multimodal AI

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Evaluation Engineer for LLMs & Multimodal Systems
AI Evaluation Engineer for LLMs & Multimodal Systems

ByteDance • San Jose (CA)

On-site
USD 120,000 - 190,000
Medical Insurance
Dental Insurance
Vision Insurance
+9
Data Infrastructure Engineer for ML Pipelines
Data Infrastructure Engineer for ML Pipelines

Luma • Redwood City (CA)

On-site
USD 150,000 - 240,000
Robotics Systems Engineer: Real-World Model Evaluation
Robotics Systems Engineer: Real-World Model Evaluation

Luma AI • United States

Remote
USD 130,000 - 190,000
Senior AI Systems Architect & Product Leader
Senior AI Systems Architect & Product Leader

Luma AI • San Francisco (CA)

On-site
USD 210,000 - 320,000
Software Engineer - Data Infrastructure
Software Engineer - Data Infrastructure

Luma • Redwood City (CA)

On-site
USD 150,000 - 240,000
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000