Applied Researcher: On-Device Multimodal Reasoning

Socket.dev

Sunnyvale (CA)

On-site

USD 200,000 - 320,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Apple is seeking an Applied Researcher to advance on-device multimodal reasoning for vision-language systems. The role focuses on compact VLM design that can reason over images, video, and 3D content within strict compute and memory budgets on Apple hardware.

You will partner with hardware, compiler, and ML teams to push the boundaries of small-model reasoning, enabling real-time multimodal understanding while preserving privacy and offline capability.

Qualifications

  • MS in CS/ML/CV or related field or equivalent practical experience.
  • Strong foundation in deep learning with experience in LLM or VLM training, post-training, or inference optimization.
  • Demonstrated experience with multimodal models in resource-constrained environments, including reasoning quality, decoding, or model compression.

Responsibilities

  • Own the reasoning component of the on-device multimodal stack, designing compact VLMs for on-device use.
  • Collaborate with hardware, compiler, and ML researchers to optimize inference under tight compute and memory budgets.
  • Research and implement efficient decoding, quantization, distillation, and KV-cache compression for real-time performance.

Skills

Python
PyTorch
Multimodal models
Model compression

Education

MS in CS/ML/CV or related field
PhD (preferred)

Tools

CoreML
MLX
TensorRT-LLM
llama.cpp
vLLM

Job description

The Video Computer Vision (VCV) organization is an applied research and engineering team developing real-time, on-device Computer Vision and Machine Perception technologies across Apple products. Within VCV, our team builds next-generation multimodal AI systems that combine on-device multimodal encoders, large language models, and foundation models to create intelligent systems capable of understanding, reasoning, and acting across language, vision, audio, and tools. Our work is deeply integrated into the Apple ecosystem, partnering across hardware, software, and ML teams to deliver real-time, scalable, and privacy-preserving experiences reaching millions of users.

Description

We are seeking an Applied Researcher with deep expertise in multimodal reasoning at small model scale — making vision-language models in the smaller regime (under ~10B parameters, down to sub-1B) think, plan, and act reliably under strict compute, memory, and latency constraints. In this role you will own the reasoning side of the on-device multimodal stack: designing compact VLMs that reason over images, video, and 3D scene content; compressing and distilling the reasoning process itself; and engineering the decoding and inference path that makes multi-step reasoning affordable on an Apple device. This role offers the unique opportunity to define what on-device intelligence looks like for hundreds of millions of users. You'll push the boundaries of what small models can achieve — enabling real-time multimodal understanding and multi-step reasoning without reliance on cloud connectivity. You'll collaborate with hardware teams, compiler engineers, and ML researchers to unlock capabilities that few organizations can deliver at Apple's scale and quality bar.This role spans multiple dimensions of efficient on-device reasoning — including VLM architecture and connector design, reasoning post-training (SFT/RL), chain-of-thought compression, speculative and structured decoding, visual token reduction, quantization and distillation, and hardware-aware inference optimization. A core focus of this role is efficient reasoning: compressed and latent chain-of-thought, reasoning distillation from frontier teachers, adaptive test-time compute (knowing when — and how long — to think), speculative and structured decoding, KV-cache compression, and visual-token efficiency. A second focus is reasoning over real-time visual perception experts. Rather than consuming pixels alone, the VLM should be able to invoke and reason over the outputs of specialist on-device vision models — feed-forward 3D scene and geometry estimators (VGGT-style reconstruction, depth, camera pose), human body and hand mesh/pose recovery, object detectors, localizers, and trackers — and fuse those structured, metric outputs into its reasoning about the scene. This raises real research questions: how to represent geometry, body parameters, and detections compactly in a token-budgeted context; how to schedule which experts run at which frame rate within a real-time budget; and how to train a small model to invoke, trust, and cross-check them. Efficiency is treated as a first-class metric here: reasoning quality is measured at a fixed latency, memory, and power budget.

Minimum Qualifications
  • MS in Computer Science, Machine Learning, AI, Computer Vision, or a related field (or equivalent practical experience)
  • Strong foundation in deep learning, with specific experience in LLM or VLM training, post-training, or inference optimization
  • Demonstrated experience working with multimodal models (vision-language models, multimodal LLMs) in resource-constrained environments, including hands-on work with reasoning quality, decoding, or model compression
  • Proficiency in Python and modern deep learning frameworks (PyTorch preferred), with familiarity with inference and optimization toolchains (quantization, distillation, pruning — e.g., vLLM/SGLang, llama.cpp, MLX, CoreML)
Preferred Qualifications
  • PhD with research in efficient multimodal reasoning, LLM reasoning, model compression, efficient inference/decoding, or lightweight VLM architectures
  • Experience training or post-training vision-language models end-to-end — connector/projector design, visual instruction tuning, resolution and token-budget trade-offs, small-model recipes
  • Hands-on experience with reasoning techniques: chain-of-thought distillation and compression, latent/implicit reasoning, reward-guided decoding, RL for reasoning (GRPO, RLVR, STaR), or test-time compute allocation
  • Expertise in decoding and serving optimizations: speculative decoding, structured/grammar-constrained generation, KV-cache quantization and eviction, continuous batching, long-context inference
  • Experience combining LLMs with real-time perception models — 3D reconstruction and geometry (VGGT, DUSt3R/MASt3R-style, SLAM, monocular depth), human pose and body/hand mesh recovery (SMPL-family), detection, segmentation, or tracking — and with spatial or 3D-grounded reasoning and embodied/spatial VQA
  • Experience deploying LLM or multimodal models on mobile or edge hardware (CoreML, MLX, TensorRT-LLM, or equivalent), with attention to ANE/GPU kernel and memory constraints
  • Experience with quantization-aware training, mixed-precision inference, and knowledge distillation for vision-language models
  • Familiarity with efficient vision encoders and self-supervised/joint-embedding pretraining (V-JEPA, I-JEPA, MAE, DINO/DINOv2, SigLIP, CLIP), Mamba/SSM vision backbones, or streaming architectures for real-time video with fixed memory budgets
  • Interest in Video-LLMs, long-video reasoning, and world models for prediction and planning
  • Publication record in top-tier venues is a plus (NeurIPS, ICML, ICLR, CVPR, ECCV, ACL, MLSys, ICRA, etc.)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied Researcher: On-Device Multimodal Reasoning
Applied Researcher: On-Device Multimodal Reasoning

Apple Inc. • Sunnyvale (CA)

Hybrid
USD 150,000 - 278,000
Applied Researcher: On-Device Multimodal Reasoning
Applied Researcher: On-Device Multimodal Reasoning

Socket.dev • Sunnyvale (CA)

On-site
USD 200,000 - 320,000
Multimodal AI Researcher
Multimodal AI Researcher

Socket.dev • Sunnyvale (CA)

Hybrid
USD 150,000 - 230,000
Machine Learning Researcher Multi-Modal Reasoning
Machine Learning Researcher Multi-Modal Reasoning

Socket.dev • Cupertino (CA)

On-site
USD 250,000 - 350,000
Multimodal LLMs Research Engineer
Multimodal LLMs Research Engineer

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
Research Manager, Multimodal Reasoning - SIML
Research Manager, Multimodal Reasoning - SIML

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 238,000 - 356,000
Relocation
Machine Learning Systems Engineer – Video Computer Vision
Machine Learning Systems Engineer – Video Computer Vision

Apple • Sunnyvale (CA)

On-site
USD 190,000 - 240,000
On-Device Multimodal Reasoning Architect
On-Device Multimodal Reasoning Architect

Apple Inc. • Sunnyvale (CA)

Hybrid
USD 150,000 - 278,000
Machine Learning Researcher Multi-Modal Reasoning
Machine Learning Researcher Multi-Modal Reasoning

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Comprehensive medical and dental
Retirement benefits
Discounted Apple products
+3
Research Manager, Multimodal Reasoning - SIML
Research Manager, Multimodal Reasoning - SIML

Apple Inc. • Seattle (WA)

On-site
USD 238,000 - 356,000
Stock programs and RSU awards
Employee Stock Purchase Plan
Education reimbursement