AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

Apple Inc.

Seattle (WA)

On-site

USD 142,300 - 263,300

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Inc. in Seattle, WA seeks an AI Research Scientist focusing on multimodal foundation models—architecture, pre-training, and distillation—for on-device intelligence. You will push novel architectures and efficient training methods that translate to real Apple experiences.

You will explore teacher–student distillation, data strategies, and energy-aware designs, with opportunities for publication and selective open-sourcing. Collaboration across hardware, software, and AI teams is required.

Qualifications

  • Experience designing, implementing, and running large-scale pre-training experiments for large language models.
  • Experience with LLM pre-training topics including architecture, objectives, data mixtures, tokenization, curricula, scaling, optimization.
  • Proficiency with PyTorch or JAX and distributed training systems.
  • Experience evaluating pretrained models on language understanding, reasoning, instruction following, or multimodal capabilities.
  • Strong understanding of transformer-based architectures and scalable training.
  • Master’s degree or equivalent practical experience in ML/CS.

Responsibilities

  • Advance architectures, pre-training methods, and distillation techniques for multimodal foundation models.
  • Experiment at scale across architecture, data strategies, and optimization.
  • Translate successful ideas into foundation-model technologies for Apple products.
  • Collaborate across model research, data, systems, and hardware boundaries.

Skills

LLM pretraining experiments
LLM pretraining topics
PyTorch/JAX
Multimodal evaluation
Transformer architectures
MS or equivalent experience

Education

Master’s degree or equivalent

Job description

AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

Seattle, Washington, United States Machine Learning and AI

The Multimodal Intelligence Team is building the next generation of foundation models for Apple experiences. We are looking for a research scientist to advance the architectures, pre‑training methods, and distillation techniques that make highly capable multimodal models practical across the Apple ecosystem. Our research spans the full foundation‑model lifecycle: model architecture, pre‑training objectives, data mixtures, optimization, scaling, distillation, and evaluation. A defining challenge of our work is to develop models that combine broad intelligence with the memory, latency, energy, and privacy requirements of on‑device deployment.

You will have the opportunity to shape new research directions, conduct ambitious experiments at scale, and translate successful ideas into foundation‑model technologies that can reach Apple products. Where appropriate, this work may also lead to publications and the open sourcing of selected models, research artifacts, evaluations, or tools.

Description

In this role, you will investigate fundamental questions about how multimodal foundation models should be designed, trained, and distilled. You will develop and evaluate new model architectures, pre‑training objectives, data strategies, optimization methods, and teacher–student learning techniques. Your work will explore how capabilities developed in large foundation models can be effectively transferred to smaller, more efficient models without treating distillation as an isolated downstream step. A major focus of the role will be the co‑development of frontier models and efficient models for Apple silicon and on‑device intelligence. This includes designing architectures that distill effectively, studying how teacher and student models should be trained together, and developing distillation methods that preserve reasoning, multimodal understanding, instruction following, and other important capabilities under constrained model capacity. Rather than treating deployment constraints as an afterthought, you will incorporate them into the research process—from early architecture experiments and pre‑training through distillation and final model evaluation.

You may thrive in this role if you:

  • Want to invent new foundation‑model architectures rather than only adapt existing models.
  • Enjoy combining scientific ambition with real compute, memory, latency, and energy constraints.
  • Believe that small and efficient models can be a frontier research problem, not merely a compression exercise.
  • Are comfortable working across model research, data, systems, and hardware boundaries.
  • Care about translating research into private, useful, and deeply integrated intelligent experiences.
  • Want your work to have both product impact and a presence in the broader research community.

Potential research directions include:

  • Novel dense, recurrent, state‑space, mixture‑of‑experts, and hybrid foundation‑model architectures.
  • Multimodal pre‑training across language, images, video, audio, and sensor‑derived representations.
  • Compute‑optimal model and data scaling, including data mixtures, curricula, tokenization, and training objectives.
  • Architecture and algorithm co‑design for memory‑efficient and energy‑efficient inference on Apple silicon.
  • Offline and on‑policy distillation using teacher‑generated data, logits, representations, rationales, and other supervision signals.
Minimum Qualifications
  • Hands‑on experience designing, implementing, and running large‑scale pre‑training experiments for large language models.
  • Experience with LLM pre‑training topics such as model architecture, training objectives, data mixtures, tokenization, curricula, scaling, and optimization.
  • Strong proficiency with modern deep learning frameworks such as PyTorch or JAX and distributed training systems.
  • Experience evaluating pre‑trained models across language understanding, reasoning, instruction following, or multimodal capabilities.
  • Strong understanding of transformer‑based architectures and current approaches to efficient or scalable foundation‑model training.
  • Master’s degree, or equivalent practical experience in machine learning, computer science, or a related technical field.
Preferred Qualifications
  • Experience contributing to major foundation‑model pre‑training efforts or leading architecture experiments that influenced a large training run.
  • Research contributions in model architecture, scaling laws, multimodal pre‑training, optimization, efficient attention, mixture‑of‑experts, state‑space models, or related areas.
  • Experience with knowledge distillation, including offline or off‑policy distillation, on‑policy distillation, self‑distillation, sequence‑level distillation, logic matching, or representation transfer.
  • Experience designing teacher–student training pipelines or transferring capabilities from large foundation models to smaller models.
  • Experience with multimodal models spanning language, vision, video, audio, or other sensor modalities.
  • Understanding of inference efficiency, memory hierarchy, hardware accelerators, or hardware–software co‑design.
  • Strong publication record, influential open‑source contributions, or an equivalent record of applied research impact.

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides an opportunity to progress as you grow and develop within a role. The base pay range for this role is between $142,300 and $263,300, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant.

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI ML Research Scientist - Foundation Models & Multimodal Intelligence
AI ML Research Scientist - Foundation Models & Multimodal Intelligence

AIToolboard • United States

On-site
USD 139,000 - 259,000
Apple employee stock programs
Discretionary RSU awards
Relocation assistance
+2
AIML - Machine Learning Engineer, Foundation Models
AIML - Machine Learning Engineer, Foundation Models

Apple Inc. • Seattle (WA)

On-site
USD 184,700 - 324,800
Medical and dental coverage
Employee stock programs
Relocation assistance
Senior Research Engineer, Training Data Infrastructure in Foundation Models
Senior Research Engineer, Training Data Infrastructure in Foundation Models

Apple Inc. • Cupertino (CA)

On-site
USD 184,700 - 324,800
Stock options
Employee stock purchase plan
Medical and dental coverage
+2
AIML - Machine Learning Researcher, Data and ML Innovation
AIML - Machine Learning Researcher, Data and ML Innovation

Apple Inc. • Santa Clara (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and free services
+1
AIML - Machine Learning Researcher - Multimodal Agent
AIML - Machine Learning Researcher - Multimodal Agent

Apple Inc. • Santa Clara (CA)

On-site
USD 184,700 - 324,800
AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation
AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 210,000
Multimodal Machine Learning Researcher
Multimodal Machine Learning Researcher

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Stock programs
Discretionary bonuses
Relocation assistance
+1
Research Manager, Multimodal Reasoning - SIML
Research Manager, Multimodal Reasoning - SIML

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 238,000 - 356,000
Relocation
AIML - Sr Machine Learning Engineer, Data and ML Innovation
AIML - Sr Machine Learning Engineer, Data and ML Innovation

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
Employee stock programs
Discretionary bonuses
Relocation assistance
+2
Multimodal AI Researcher
Multimodal AI Researcher

Apple Inc. • Sunnyvale (CA)

Hybrid
USD 150,000 - 278,000
Stock programs
Discretionary bonuses
Relocation
+1