Vision-Language ML Architect & Inventor of Novel VLMs

Arcade

Presidio (TX)

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Daily catered lunch
Company events
Equal opportunity employer

Job summary

Arcade is seeking a high-caliber Vision-Language Model expert to lead innovative research and develop production-ready AI systems that evaluate consumer product imagery with advanced multimodal capabilities.

You will invent architectures and training paradigms, pushing beyond standard models to quantify aesthetics, market value, and context-aware captions for on-demand manufacturing platforms. Join a world-class team in San Francisco's Presidio.

Qualifications

  • Ph.D. in CS/AI/ML or closely related field.
  • Strong publication record in top AI/CV/NLP venues.
  • Deep theoretical and practical understanding of Vision-Language Models and multimodal architectures.
  • Proficient in Python and PyTorch with ability to write optimized training loops.

Responsibilities

  • Pioneer novel neural architectures, loss functions, and multimodal integration techniques.
  • Lead fine-tuning and adaptation of state-of-the-art Vision-Language Models for complex tasks.
  • Develop foundations to extract nuanced signals from product images and translate them into learnable objectives.
  • Guide dataset creation, curation, and augmentation for domain-specific concepts.
  • Invent evaluation frameworks and custom metrics for abstract tasks without standard benchmarks.
  • Collaborate with engineering to translate research into scalable, production-ready systems.

Skills

VLM expertise
Multimodal architectures
Transformer models
Python
Problem solving

Education

Ph.D. in Computer Science/AI/ML

Tools

PyTorch

Job description

Arcade is seeking a high-caliber Vision-Language Model expert to lead innovative research and develop production-ready AI systems that evaluate consumer product imagery with advanced multimodal capabilities.

You will invent architectures and training paradigms, pushing beyond standard models to quantify aesthetics, market value, and context-aware captions for on-demand manufacturing platforms. Join a world-class team in San Francisco's Presidio.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Vision-Language AI Scientist: Invent Multimodal Systems
Vision-Language AI Scientist: Invent Multimodal Systems

Arcade • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Daily catered lunch
Company events
Research Scientist
Research Scientist

Arcade • Presidio (TX)

On-site
USD 180,000 - 240,000
Daily catered lunch
Company events
Equal opportunity employer
Research Scientist
Research Scientist

Arcade • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Daily catered lunch
Company events
Lead Vision-Language Model Post-Training Engineer
Lead Vision-Language Model Post-Training Engineer

Liquid AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive salary with equity
100% health premiums for employees and dependents
401(k) matching
+1
Founding Computer Vision Engineer – Multimodal AI & VLMs
Founding Computer Vision Engineer – Multimodal AI & VLMs

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Vision AI Engineer for Multimodal LLMs (Equity Eligible)
Vision AI Engineer for Multimodal LLMs (Equity Eligible)

Pinterest • San Francisco (CA)

Hybrid
USD 139,000 - 286,000
Equity
Vision-Language Research Engineer
Vision-Language Research Engineer

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Senior ML Engineer - Vision & Multimodal (Production)
Senior ML Engineer - Vision & Multimodal (Production)

Clearview AI • United States

On-site
USD 150,000 - 200,000
Medical, Dental, Vision
STD and LTD Plans
ML Engineer – Vision-Language for Motion & Autonomy
ML Engineer – Vision-Language for Motion & Autonomy

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
Vision-Language Models (VLMs)
Vision-Language Models (VLMs)

TalentOla • Waukesha (WI)

On-site
USD 120,000 - 150,000