Multimodal AI Researcher

Socket.dev

Sunnyvale (CA)

Hybrid

USD 150,000 - 230,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple in Sunnyvale is seeking a Multimodal AI Researcher to push the boundaries of foundation models for real-time multimodal data, including video, audio, and text. You will work on interactive models, audio-to-audio modeling, and streaming multimodal systems, driving data requirements, validation strategies, and delivering research that informs product features.

Ideal candidates have a BS with 3+ years of experience, hands-on work with LLMs and VLMs, and strong Python/PyTorch skills.

Qualifications

  • BS and a minimum of 3 years relevant industry experience.
  • Experience building models for multimodal perception systems.
  • Experience working with LLMs and VLMs.
  • Software engineering skills with Python and PyTorch.
  • Curiosity and willingness to learn new things to improve solutions.

Responsibilities

  • Conduct algorithm research and development for multimodal foundational models and agents.
  • Collaborate with data science, ML, and product teams to define requirements and KPIs.
  • Translate research into product features used by millions of users across Apple devices.

Skills

Multimodal AI
Python
PyTorch
LLMs
VLMs
Research experience

Education

BS in related field
MS or PhD in related fields

Job description

The Video Computer Vision organization is working on breakthrough technologies for future Apple products. Our team delivers cutting-edge AI, machine learning, computer vision and graphics algorithms that power technologies including human understanding, perception, digital humans, multimodal generative AI, and agents. Our algorithms ship across a range of Apple products, including iPhone and Apple Vision Pro, where our work has contributed to technologies like Personalized Spatial Audio, EyeSight, and Persona as well as future Apple products. We are an applied research group, we push the state of the art and then bring it to product. In this role, you will collaborate with world-class experts in AI, ML, Software, and Hardware to tackle fundamental challenges in human-centric solutions that will impact millions of users across Apple's ecosystem.

Description

We are looking for a Multimodal AI Researcher with a strong background in developing foundation models for generative AI and multimodal systems that integrate various types of real-time sensor data such as video and audio with other modalities like text. Our ongoing investigations include interactive models, audio-to-audio modeling and systems. You will work on hard, open research problems in multimodal generative AI and agents, and you will see that work through to real features used by millions of people. You will collaborate with others to drive data requirements, validation strategies, and key performance indicators, and conduct algorithm research and development that serves product needs. We hire researchers who are highly motivated and deeply care about shipping. A successful candidate will stay up-to-date with the latest advancements in multimodal foundations models and applying this knowledge to drive innovation, but also take a practical approach to problem solving and software engineering.

Minimum Qualifications
  • BS and a minimum of 3 years relevant industry experience.
  • Experience building models for multimodal perception systems.
  • Experience working with LLMs and VLMs.
  • Software engineering skills and proficiency in Python and PyTorch.
  • Curiosity and willingness to learn new things in order to improve the quality of their solutions.
Preferred Qualifications
  • MS or PhD in computer vision, computer graphics, machine learning, computer science, computer engineering or related fields.
  • Experience in developing, training/tuning foundation models and multimodal LLMs.
  • Experience with training and troubleshooting generative architectures such as diffusion, reinforcement learning, flow matching or normalizing flow at scale.
  • Experience with real-time or streaming multimodal models.
  • Experience with speech understanding and generation.
  • Experience applying reinforcement learning to help post-train foundation models.
  • Excellent communication and experience working with multi-functional teams.
  • Self-motivated with proven track record to optimally prioritize and deliver tasks on schedule.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Multimodal AI Researcher
Multimodal AI Researcher

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
Medical and dental coverage
Retirement benefits
Employee stock purchase plan
+2
Multimodal AI Researcher: Generative Models, Realtime Vision
Multimodal AI Researcher: Generative Models, Realtime Vision

Socket.dev • Sunnyvale (CA)

Hybrid
USD 150,000 - 230,000
Multimodal AI Researcher: Generative Models & Vision
Multimodal AI Researcher: Generative Models & Vision

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
Medical and dental coverage
Retirement benefits
Employee stock purchase plan
+2
Multimodal LLMs Research Engineer
Multimodal LLMs Research Engineer

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
AI/ML Software Engineer
AI/ML Software Engineer

Socket.dev • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
AIML - Machine Learning Researcher - Multimodal Agent
AIML - Machine Learning Researcher - Multimodal Agent

Apple Inc. • Santa Clara (CA)

On-site
USD 184,700 - 324,800
Applied Researcher: On-Device Multimodal Reasoning
Applied Researcher: On-Device Multimodal Reasoning

Socket.dev • Sunnyvale (CA)

On-site
USD 200,000 - 320,000
Machine Learning Systems Engineer – Video Computer Vision
Machine Learning Systems Engineer – Video Computer Vision

Apple • Sunnyvale (CA)

On-site
USD 190,000 - 240,000
Research Engineer, Multimodal
Research Engineer, Multimodal

character • Redwood City (CA)

On-site
USD 100,000 - 150,000
AIML - Machine Learning Researcher, DMLI- Image/Video Generation
AIML - Machine Learning Researcher, DMLI- Image/Video Generation

Apple Inc. • Seattle (WA)

On-site
USD 184,700 - 324,800
Medical & Dental
Retirement benefits
Employee stock plan
+1