Research Scientist

Arcade

Presidio (TX)

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Daily catered lunch
Company events
Equal opportunity employer

Job summary

Arcade is seeking a high-caliber Vision-Language Model expert to lead innovative research and develop production-ready AI systems that evaluate consumer product imagery with advanced multimodal capabilities.

You will invent architectures and training paradigms, pushing beyond standard models to quantify aesthetics, market value, and context-aware captions for on-demand manufacturing platforms. Join a world-class team in San Francisco's Presidio.

Qualifications

  • Ph.D. in CS/AI/ML or closely related field.
  • Strong publication record in top AI/CV/NLP venues.
  • Deep theoretical and practical understanding of Vision-Language Models and multimodal architectures.
  • Proficient in Python and PyTorch with ability to write optimized training loops.

Responsibilities

  • Pioneer novel neural architectures, loss functions, and multimodal integration techniques.
  • Lead fine-tuning and adaptation of state-of-the-art Vision-Language Models for complex tasks.
  • Develop foundations to extract nuanced signals from product images and translate them into learnable objectives.
  • Guide dataset creation, curation, and augmentation for domain-specific concepts.
  • Invent evaluation frameworks and custom metrics for abstract tasks without standard benchmarks.
  • Collaborate with engineering to translate research into scalable, production-ready systems.

Skills

VLM expertise
Multimodal architectures
Transformer models
Python
Problem solving

Education

Ph.D. in Computer Science/AI/ML

Tools

PyTorch

Job description

About Arcade

Arcade is building the world's first AI physical product creation platform, where imagination becomes reality. Our platform lets anyone design, purchase, and sell custom, manufacturable products using natural language and generative AI. We believe everyone should have the power to create physical goods as easily as they post online, and we're building the infrastructure to make that real for both consumers and businesses.

We've raised $42M from a world-class group of investors, including Reid Hoffman, Forerunner Ventures (Kirsten Green), Canaan Partners (Laura Chau), Adverb Ventures (April Underwood), Factorial Funds (Sol Bier), Offline Ventures (Brit Morin), Sound Ventures (Ashton Kutcher), Inspired Capital (Alexa von Tobel), and Torch Capital (Jonathan Keidan). Our angel investors include Elad Gil, Ev Williams, Marissa Mayer, Sara Beykpour, Kayvon Beykpour, Anna Veronika Dorogush, Eugenia Kuyda, David Luan, Sharon Zhou, Kelly Wearstler, Karlie Kloss, Colin Kaepernick, Christy Turlington Burns, and Jeff Wilke.

Arcade is headquartered in San Francisco's Presidio and led by serial entrepreneur Mariam Naficy (Minted, Eve), and a mission-driven team from Google, Apple, Stability AI, Glean, NVIDIA, Databricks, LinkedIn, Stanford, MIT, Berkeley, and more. Arcade's Chief AI Officer is Varun Jampani, a leading researcher who co-authored Dreambooth and created Stable Diffusion 3.5, among other things. Raghudeep Gadde, Head of Research at Arcade, was formerly a Principal Scientist at Amazon. Together, we're pioneering a new category at the intersection of AI, personal expression, and on-demand manufacturing, and we're building fast.

The Role

We are seeking a high-caliber, deeply innovative Vision-Language Model (VLM) expert to lead our efforts in teaching foundation models how to evaluate consumer products like an expert appraiser.

This is not a standard implementation role. You will be expected to invent new technologies, design novel architectures, and author proprietary training paradigms when existing open-source or commercial models fall short. You will go beyond simple object detection, engineering systems that can estimate highly abstract and valuable aspects of product images—such as generating rich, context-aware captions, predicting precise market price points, and evaluating subjective aesthetic quality or "beauty" scores.

If you are a pioneer who thrives on solving unsolved multimodal problems and wants your inventions to power a groundbreaking production platform, we want to hear from you.

Responsibilities
  • Innovation & Invention: Pioneer novel neural architectures, loss functions, and multimodal integration techniques. We expect you to invent new AI technologies and methodologies to solve unprecedented challenges in visual product perception.

  • VLM Fine-Tuning & Adaptation: Lead the deep fine-tuning, adaptation, and structural optimization of state-of-the-art Vision-Language Models (e.g., Qwen-VLM, PaliGemma2 etc.) for targeted, highly complex computer vision tasks.

  • Complex Attribute Estimation: Develop the mathematical and architectural foundations to extract highly nuanced signals from product images, translating subjective concepts (like aesthetics and market value) into rigorous, learnable objectives.

  • Dataset Strategy: Guide the creation, curation, and algorithmic augmentation of specialized multimodal datasets required to teach foundation models novel, domain-specific concepts.

  • Evaluation & Metrics: Invent rigorous evaluation frameworks and custom metrics to accurately measure model performance on abstract tasks where standard academic benchmarks do not exist.

  • Cross-Functional Collaboration: Work closely with engineering teams to ensure your proprietary research and new technologies translate seamlessly into scalable, production-ready systems.

Qualifications
  • Education: Ph.D. in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a strictly related field.

  • Publication Record: A strong track record of advancing the state-of-the-art, evidenced by first-author publications in top-tier AI, CV, or NLP venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, ACL).

  • Domain Expertise: Deep theoretical and practical understanding of Vision-Language Models, multimodal architectures, and modern transformer-based computer vision. You must understand the math and mechanics under the hood.

  • Technical Stack: Expert-level proficiency in Python and deep learning frameworks (specifically PyTorch), with the ability to write optimized training loops when necessary.

  • Problem Solving: A proven track record of formulating highly ambiguous, real-world visual problems into rigorous, solvable machine learning tasks.

  • Industry Experience (bonus): Prior industrial experience as a Research Scientist or Machine Learning Engineer, specifically involving the deployment of deep learning models to large-scale production environments.

  • E-commerce/Product ML (bonus): Previous experience applying machine learning to product imagery, retail technology, or computational aesthetics.

Additional

Competitive compensation

Daily catered lunch prepared by our chef

Company events

Arcade is an equal opportunity employer. We're committed to building a diverse, inclusive, and supportive team, and to creating a platform where anyone, anywhere, can make something meaningful.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist
Research Scientist

Arcade • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Daily catered lunch
Company events
Staff Applied AI Engineer
Staff Applied AI Engineer

Arcade • Presidio (TX)

On-site
USD 180,000 - 240,000
Lunch provided daily
Company events
Competitive compensation
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Arcade • San Francisco (CA)

On-site
USD 180,000 - 300,000
Unlimited PTO
Rich health/401(k) plans
Meeting-light culture
+2
Applied AI Engineer
Applied AI Engineer

arcade-ai • Tracy (CA)

On-site
USD 90,000 - 120,000
Competitive compensation
Lunch provided daily
Company events
Senior Applied AI Engineer
Senior Applied AI Engineer

Cerebras • San Francisco (CA)

Hybrid
USD 130,000 - 265,000
Unlimited PTO
Rich health/401(k) plans
Meeting-light culture
+2
Vision-Language ML Architect & Inventor of Novel VLMs
Vision-Language ML Architect & Inventor of Novel VLMs

Arcade • Presidio (TX)

On-site
USD 180,000 - 240,000
Daily catered lunch
Company events
Equal opportunity employer
Research Engineer, Visual Knowledge Work
Research Engineer, Visual Knowledge Work

Anthropic • New York (NY)

Hybrid
USD 350,000 - 850,000
Generous vacation
Parental leave
Flexible working hours
Vision-Language AI Scientist: Invent Multimodal Systems
Vision-Language AI Scientist: Invent Multimodal Systems

Arcade • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Daily catered lunch
Company events
Research Engineer, Multimodal Data
Research Engineer, Multimodal Data

Eventual • San Francisco (CA)

On-site
USD 120,000 - 150,000
Catered lunches and dinners
Commuter benefit
Health, vision, and dental coverage
+2
ML Engineer
ML Engineer

RiseMe • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Equity compensation
Medical insurance
Flexible PTO
+8