Multimodal 3D Vision & LLM Researcher

Javelin Venture Partners

Greater London

On-site

GBP 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Niantic Spatial in London is seeking a Computer Vision Researcher with a strong background in Large Language Models to bridge 3D vision and language. You will work on aligning spatial geometry with language to enable context-aware navigation and open-ended reasoning about real-world environments.

The role involves leading semantic grounding research, developing continuous semantic capabilities for evolving 3D maps, and mentoring the London R&D hub team.

Qualifications

  • PhD (or equivalent) in Computer Vision, ML or Robotics with a focus on Multimodal/Semantic understanding.
  • 4+ years of ML research experience with a track record of shipping models bridging 3D vision and language.
  • Expert knowledge of 3D geometry (SfM, SLAM, VPS) and Transformer-based architectures (VLMs).
  • First-author publications at top-tier venues (CVPR, NeurIPS, ICLR) focusing on VLMs, scene understanding or semantic segmentation.

Responsibilities

  • Architect Semantic Grounding: lead research into cross-modal grounding connecting 3D spatial features with language embeddings.
  • Scale Understand Capabilities: develop algorithms for continuous semantics for evolving 3D maps.
  • Agentic Frameworks: build the spatial brain for Embodied AI enabling robots and drones to perform mission-level reasoning.
  • Multimodal Benchmarking: define standards for spatial common sense in LLMs; create evaluations for 3D scene interpretation.

Skills

3D geometry
Transformer-based architectures
Production-quality code (PyTorch/JAX)

Education

PhD in Computer Vision / ML / Robotics

Tools

PyTorch
JAX
ROS

Job description

Niantic Spatial in London is seeking a Computer Vision Researcher with a strong background in Large Language Models to bridge 3D vision and language. You will work on aligning spatial geometry with language to enable context-aware navigation and open-ended reasoning about real-world environments.

The role involves leading semantic grounding research, developing continuous semantic capabilities for evolving 3D maps, and mentoring the London R&D hub team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Computer Vision Researcher (VLM)
Computer Vision Researcher (VLM)

Niantic Spatial, Inc. • Greater London

On-site
GBP 70,000 - 90,000
Senior CV Researcher: Multimodal Spatial AI (LLM)
Senior CV Researcher: Multimodal Spatial AI (LLM)

Niantic Spatial, Inc. • Greater London

Hybrid
GBP 70,000 - 90,000
Computer Vision Researcher (VLM)
Computer Vision Researcher (VLM)

Javelin Venture Partners • Greater London

On-site
GBP 120,000 - 180,000
Multimodal Vision Research Scientist
Multimodal Vision Research Scientist

IC Resources • Greater London

On-site
GBP 70,000 - 120,000
Multimodal Vision AI Scientist: Research & Impact in London
Multimodal Vision AI Scientist: Research & Impact in London

IC Resources Recruitment • Greater London

On-site
GBP 70,000 - 110,000
Computer Vision Engineer
Computer Vision Engineer

microTECH Global LTD • Greater London

On-site
GBP 70,000 - 110,000
Multimodal Vision Engineer — AI Systems & Research
Multimodal Vision Engineer — AI Systems & Research

microTECH Global LTD • Greater London

On-site
GBP 70,000 - 110,000
3D Reconstruction Scientist: SfM, SLAM & Dense Geometry
3D Reconstruction Scientist: SfM, SLAM & Dense Geometry

SpAItial AI • Greater London

On-site
GBP 90,000 - 120,000
Model Research Scientist - Agentic AI & Multimodal (Hybrid)
Model Research Scientist - Agentic AI & Multimodal (Hybrid)

H Company • Greater London

Hybrid
GBP 120,000 - 180,000
Competitive salary
Opportunities for professional growth
Computer Vision Research Scientist
Computer Vision Research Scientist

IC Resources • Greater London

On-site
GBP 70,000 - 120,000