Research Scientist - VLM

Storm3

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
Competitive salary

Job summary

Storm3 in San Francisco is seeking researchers with expertise in VLMs, vision-language, and multimodal understanding to advance the capabilities of our LLMs.

PhD in CS/ML is preferred, with 2+ years of industry experience in VLMs and a strong publication record. You will help build the multimodal data infrastructure, curate datasets, and publish breakthroughs with the team.

Qualifications

  • PhD in CS/ML/Maths preferred.
  • 2+ years industry experience focused on VLMs, Visual Understanding, 2D Image/Video.
  • Strong publication record in top conferences or contributions to leading AI models in VLMs.

Responsibilities

  • Research & develop methods in image/video understanding, vision-language, and related areas.
  • Contribute to large-scale vision data infrastructure and model training.
  • Source, process and curate high quality image and video data.
  • Publish breakthrough findings and reports in top conferences.

Education

PhD in Comp Sci/ML/Maths

Job description

Come join one of the only research institutions globally with resources to compete with top AI companies => 10s of 1000s of GPUs and Tier1 talent.

This lab is a playground for state-of-the-art research in LLMs, Agents and World Models.

Hiring for those experienced in VLMs, vision-language, image/video understanding & reasoning to contribute to the multimodal capabilities of their LLMs.

Responsibilities:
  • Research & develop novel methods in image/video understanding & recognition, document understanding, charts & figures, and computer-use or GUI
  • Contribute towards large-scale vision data infrastructure and model training
  • Source, process and curate high quality image and video data
  • Publish breakthrough findings and reports in top conferences
Requirements:
  • PhD in Comp Sci/ML/Maths preferred
  • Strong publication record in top conferences, or contribution to leading AI models - particularly in VLMs, visual understanding, world models or physical AI
  • 2+ years industry experience focused on VLMs, Visual Understanding, 2D Image/Video
  • Strong in navigating ambiguity and impacting research direction
Why apply:
  • Opportunity to join a fast-growing core team that are already pushing AI breakthroughs
  • Highly competitive salary package
  • Work alongside ambitious and bright superstars from tech and academia
  • Medical, Dental and Vision Insurance
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Visual Knowledge Work
Research Engineer, Visual Knowledge Work

Anthropic • New York (NY)

Hybrid
USD 350,000 - 850,000
Generous vacation
Parental leave
Flexible working hours
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Clearview AI • United States

On-site
USD 150,000 - 200,000
Medical, Dental, Vision
STD and LTD Plans
Research Engineer - Vision Language Models / Multimodal AI / Computer Vision
Research Engineer - Vision Language Models / Multimodal AI / Computer Vision

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, Vision / Language
Member of Technical Staff, Vision / Language

xdof.ai • San Mateo (CA)

On-site
USD 120,000 - 160,000
Competitive compensation and equity
Comprehensive health and wellness benefits
Collaborative and fast-paced work environment
Research Manager
Research Manager

Harnham • United States

Hybrid
USD 180,000 - 280,000
Research Scientist – World Modeling, Data
Research Scientist – World Modeling, Data

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 400,000
Medical benefits
Dental/vision
401K
+5
Research Engineer
Research Engineer

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 140,000 - 190,000
Applied Scientist III - VLM R&D
Applied Scientist III - VLM R&D

Wyze • Kirkland (WA)

On-site
USD 134,000 - 181,000
Research Engineer, Visual Knowledge Work
Research Engineer, Visual Knowledge Work

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Research Scientist - Vision Data Infrastructure
Research Scientist - Vision Data Infrastructure

Storm3 • San Francisco (CA)

On-site
USD 250,000 - 600,000
Medical, Dental and Vision Insurance
Highly competitive salary package