Research Scientist / Engineer — Multimodal Agent

lumalabs-ai

San Francisco (CA)

On-site

USD 250,000 - 450,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Luma AI seeks to define the future of multimodal AI by building and training large-scale multimodal models. You will join a team that bridges research with shipped products, working end-to-end on novel problems with no existing playbook.

This role blends the science and engineering aspects of research, focusing on modeling, data, systems, and evaluation to advance multimodal agents across big datasets and GPU clusters.

Qualifications

  • Strong foundation in machine learning, foundation models and agentic systems.
  • Deep understanding of agentic systems and approaches in LLM/VLM reasoning, coding models, LLM/VLM tool calling.
  • Hands-on experience with PyTorch and large-scale training (distributed, mixed precision, large datasets).

Responsibilities

  • Architect large-scale multimodal agentic models that use reasoning, planning, coding, and tool calling to achieve complex, multi-step multimodal work.
  • Data: Hillclimbing existing tasks and formulating new tasks through data. Design, implement, and run robust data pipelines for constructing, enriching, and filtering massive pixel datasets.
  • Systems: Train large-scale multimodal models on massive datasets and GPU clusters.
  • Evaluation: Define and build novel evaluation frameworks to measure multimodal agents.

Skills

Foundation models
Agentic systems
LLM/VLM reasoning
Large-scale training

Tools

PyTorch

Job description

About Luma AI:

Luma's mission is to build multimodal AGI. Through our research on video, 3D, and now multimodal models at Luma, we believe that AI needs to be jointly trained over all signal modalities - text, video, audio, images - analogous to the human brain.


To advance our mission, we build and operate the full stack end-to-end, spanning foundation models, inference systems, and products. This integrated approach powers technologies like Ray3, which is seeing rapidly growing adoption among Fortune 500 companies across media, entertainment, and advertising. Backed by a recent $900M Series C and our partnership with Humain to build a 2 GW compute supercluster (Project Halo), our models and the Dream Machine platform are now enabling creatives worldwide to tell some of the most impactful stories of our time.


Where You Come In:

This is a rare and foundational opportunity to define the future of multimodal AI. You will be at the forefront of building and training large-scale multimodal models, directly impacting how users interact with pixels. This role offers the chance to bridge cutting-edge research with magical, shipped products, working end-to-end on novel problems with no existing playbook.


What You'll Do:

This opportunity involves both the \"science\" and \"engineering\" parts of research, two aspects that are of equal importance.


This is a multi-stack opportunity where you will work on the intersection of modeling, data, systems, and evaluation.



  • Modeling: Architect large-scale multimodal agentic models that use reasoning, planning, coding, and tool calling to achieve complex, multi-step multimodal work.

  • Data: Hillclimbing existing tasks and formulating new tasks through data. Design, implement, and run robust data pipelines for constructing, enriching, and filtering massive pixel datasets.

  • Systems: Train large-scale multimodal models on massive datasets and GPU clusters.

  • Evaluation: Define and build novel evaluation frameworks to measure multimodal agents.


Who You Are:


  • Strong foundation in machine learning, foundation models and agentic systems.

  • Deep understanding of agentic systems and approaches in LLM/VLM reasoning, coding models, LLM/VLM tool calling.

  • Hands-on experience with PyTorch and large-scale training (distributed, mixed precision, large datasets).


What Sets You Apart (Bonus Points):

Experience in the following around data, modeling, or evaluation:



  • State-of-the-art foundation models in reasoning

  • State-of-the-art foundation models in coding

  • State-of-the-art foundation models in tool calling

  • State-of-the-art multimodal agents


Your application are reviewed by real people.


Compensation

The base pay range for this role is $250,000 - $450,000 per year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied Research Scientist / Engineer
Applied Research Scientist / Engineer

lumalabs-ai • New York (NY)

On-site
USD 200,000 - 450,000
Research Scientist / Engineer – Foundation Model: Core Research
Research Scientist / Engineer – Foundation Model: Core Research

lumalabs-ai • San Francisco (CA)

On-site
USD 250,000 - 450,000
Research Engineer - Evaluations
Research Engineer - Evaluations

lumalabs-ai • New York (NY)

On-site
USD 190,000 - 375,000
Staff Product Software Engineer
Staff Product Software Engineer

lumalabs-ai • San Francisco (CA)

On-site
USD 230,000 - 360,000
Research Scientist / Engineer - Controllability, Personalization & Productization
Research Scientist / Engineer - Controllability, Personalization & Productization

lumalabs-ai • New York (NY)

On-site
USD 200,000 - 450,000
Research Engineer - Evaluations
Research Engineer - Evaluations

Luma AI • San Francisco (CA), New York (NY)

On-site
USD 170,000 - 210,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000
Software Engineer - Product
Software Engineer - Product

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Research Scientist / Engineer – Performance Optimization
Research Scientist / Engineer – Performance Optimization

lumalabs-ai • San Francisco (CA)

On-site
USD 237,000 - 395,000