Multimodal Machine Learning Researcher

Socket.dev

Cupertino (CA)

On-site

USD 250,000 - 380,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking a senior technical leader to architect and deploy production-scale multimodal ML within the Scene Understanding domain. You will oversee training of 2D/3D vision-language models on distributed backends and guide on-device deployment of compact architectures, while shaping private, user-centric learning policies.

The role requires cross-functional collaboration with ML researchers, software engineers, and hardware/design teams to push the boundaries of on-device experiences and

Qualifications

  • MS or PhD in CS or a related field or equivalent experience.
  • Hands-on experience training LLMs or adapting LLMs for downstream tasks.
  • Experience at the intersection of NLP and vision for multimodal models.
  • Proficiency with ML toolkits (e.g., PyTorch) and Python.

Responsibilities

  • Train large-scale multimodal models on distributed backends.
  • Deploy compact neural architectures efficiently on device.
  • Learn privacy-preserving policies personalized to users.
  • Collaborate with ML researchers, software engineers, and hardware/design teams.
  • Advance on-device scene understanding and generation capabilities.

Skills

LLM training
NLP & vision
Python
PyTorch
Distributed training
Multimodal ML

Education

MS/PhD in CS or related field

Tools

PyTorch
C/C++
ObjC

Job description

Do you believe generative models can transform creative workflows and smart assistants used by billions? Do you believe it can fundamentally shift how people interact with devices and communicate? Our Scene Understanding strives to turn cutting edge research into compelling user experiences that realize all these goals and more, working on Apple Intelligence technologies such as Image Playground, Genmoji, Generative Memories, Semantic Search, and many more. We are looking for senior technical leaders experienced in architecting and deploying production scale multimodal ML. An ideal candidate has the ability to lead diverse cross functional efforts ranging from ML modeling, prototyping, validation and private learning. Solid ML fundamentals and an ability to place research contributions with respect to state of the art would be an essential part of the role. Experience with training and adapting large language models would be an important need. We are the Intelligence System Experience (ISE) team within Apple’s software organization. The team works at the intersection between multimodal machine learning and system experiences. For example, experiences like Spotlight Search, Photos Memories, Generative Playgrounds, Stickers, Smart wallpapers, etc are all areas that the team has had a significant part in delivering through ML core technologies. These experiences that our users enjoy are backed by production ML workflows, which our team works to scale through distributed training. Additionally, our team also focuses on approaches to optimizing and adapting LLMs to best suit on-device user experiences. SELECTED REFERENCES TO OUR TEAM’S WORK:

  • - https://machinelearning.apple.com/research/introducing-apple-foundation-models (https://machinelearning.apple.com/research/introducing-apple-foundation-models)
  • - https://machinelearning.apple.com/research/stable-diffusion-coreml-apple-silicon (https://machinelearning.apple.com/research/stable-diffusion-coreml-apple-silicon)
  • - https://machinelearning.apple.com/research/on-device-scene-analysis (https://machinelearning.apple.com/research/on-device-scene-analysis)
  • - https://machinelearning.apple.com/research/panoptic-segmentation (https://machinelearning.apple.com/research/panoptic-segmentation)
Description

We are looking for a candidate with a proven track record in applied ML research. Responsibilities in the role will include training large scale multimodal (2D/3D vision-language) models on distributed backends, deployment of compact neural architectures efficiently on device, and learning policies that can be personalized to the user in a privacy preserving manner. Ensuring quality in the wild, with an emphasis on fairness and model robustness would constitute an important part of the role. You will be interacting very closely with a variety of ML researchers, software engineers, hardware & design teams cross functionally. The primary responsibilities of the role would center on enriching multimodal capabilities of large language models. The user experience initiative would focus on aligning image/video content to the space of LMs for visual actions & multi-turn interactions.

Minimum Qualifications
  • M.S. or PhD in Computer Science or a related field such as Electrical Engineering, Robotics, Statistics, Applied Mathematics, or equivalent experience.
  • Hands on experience training LLMs/adapting pre-trained LLMs for downstream tasks & alignment
  • Modeling experience at the intersection of NLP and vision
  • Proficiency in ML toolkit of choice, e.g., PyTorch
  • Strong programming skills in Python
Preferred Qualifications
  • Familiarity with distributed trainingStrong programming skills in C/C++ or ObjC
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Multimodal Machine Learning Researcher
Multimodal Machine Learning Researcher

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Stock programs
Discretionary bonuses
Relocation assistance
+1
Multimodal AI Researcher
Multimodal AI Researcher

Socket.dev • Sunnyvale (CA)

Hybrid
USD 150,000 - 230,000
Applied AI Scientist - Multimodal Intelligence
Applied AI Scientist - Multimodal Intelligence

Apple • Seattle (WA)

On-site
USD 180,000 - 280,000
Multimodal LLMs Research Engineer
Multimodal LLMs Research Engineer

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
AIML - Machine Learning Researcher - Multimodal Agent
AIML - Machine Learning Researcher - Multimodal Agent

Apple Inc. • Santa Clara (CA)

On-site
USD 184,700 - 324,800
Research Manager, Multimodal Reasoning - SIML
Research Manager, Multimodal Reasoning - SIML

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 238,000 - 356,000
Relocation
Senior Multimodal ML Research Lead for On-Device AI
Senior Multimodal ML Research Lead for On-Device AI

Socket.dev • Cupertino (CA)

On-site
USD 250,000 - 380,000
Machine Learning Manager, Data for Foundation Models - SIML
Machine Learning Manager, Data for Foundation Models - SIML

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 198,300 - 342,800
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+2
Machine Learning Research Engineer, SIML - ISE
Machine Learning Research Engineer, SIML - ISE

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation
AI Research Scientist: Multimodal Foundation Models - Architecture, Pre-Training & Distillation

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 210,000