Applied Scientist III - VLM R&D

Wyze

Kirkland (WA)

On-site

USD 134,000 - 181,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Wyze seeks an Applied Scientist to drive in-house vision-language models for smart home video understanding. You will translate breakthrough research into production VLM development and shape the next generation of physical AI at scale.

Work with tens of millions of authorized videos to uncover event patterns and inform product direction, with opportunities to publish and share open-source models and datasets with the community.

Qualifications

  • Hands-on experience training multimodal vision-language models.
  • Proficiency with Python and PyTorch.
  • Experience with visual transformers or physical world foundation models.
  • Strong research sense to identify impactful directions and translate them to practice.

Responsibilities

  • Follow the latest breakthroughs in multimodal research and apply them to Wyze VLM development.
  • Train, fine-tune, and evaluate vision-language models on large-scale real-world home video data.
  • Design rigorous evaluation pipelines for video understanding tasks like event detection and temporal reasoning.
  • Analyze user event patterns across millions of videos to guide model design and product direction.
  • Contribute to architecture and training of a physical smart home foundation model with advances in embodied AI.
  • Build rapid proofs of concept and move promising ideas to validated prototypes.
  • Publish research at top venues and contribute open-source models and datasets.
  • Help define research problems and set technical direction for industry-relevant solutions.

Skills

Hands-on experience with multimodal V×

Education

PhD in Computer Vision or ML

Job description

Overview

At Wyze, we make smart home technology accessible to everyone. We are known for disrupting markets with high-quality, affordable products—from cameras to lighting to sensors and more. We believe technology should simplify life, not complicate it. We’re a fast-moving, customer-obsessed team driven by curiosity and powered by data.

The Opportunity

We are looking for an Applied Scientist to drive the development of our in-house vision-language models (VLMs) for smart home video understanding. In this role, you will track breakthrough research from academia and the broader AI community, and rapidly translate it into our production VLM development. You will help build the next generation of smart home physical AI — foundation models that understand the physical world of the home.

You will work with tens of millions of authorized videos to deeply investigate user event patterns and build a physical smart home foundation model that impacts over 10 million Wyze households. We believe in advancing the field, not just our product: we encourage publishing your research and releasing open-source models and datasets to benefit the broader community. This is a rare opportunity to shape a category-defining product at the intersection of frontier multimodal research and real-world deployment at massive scale.

What You’ll Do
  • Follow the latest breakthroughs in multimodal and vision-language research from academia and industry, evaluate their relevance, and apply them to our in-house VLM development for smart home video understanding
  • Train, fine-tune, and evaluate multimodal vision-language models on large-scale, real-world home video data
  • Design and run rigorous evaluation pipelines to measure model quality on video understanding tasks such as event detection, activity recognition, and temporal reasoning
  • Investigate user event patterns across tens of millions of authorized videos to inform model design and product direction
  • Contribute to the architecture and training of a physical smart home foundation model, drawing on advances in visual transformers, physical world foundation models, and embodied AI
  • Build rapid proofs of concept using AI-assisted research and development workflows, and carry promising directions from idea to validated prototype
  • Publish research at top venues and contribute open-source models and datasets that help advance the community
  • Help define research problems, set technical direction, and anticipate where academic research and industry solutions are heading
What We’re Looking For
  • PhD in Computer Vision, Machine Learning, or a related field; or a Master’s degree with a strong track record of research or applied impact (publications, open-source contributions, or shipped ML systems)
  • Hands-on experience training and evaluating multimodal vision-language models
  • Experience in one or more of: visual transformer algorithm innovation, physical world foundation models, or embodied AI
  • Strong research sense: the ability to define the right problems, choose promising directions, and predict how research trends will translate into industry solutions
  • Proficiency with AI-assisted research and fast POC development — you use modern AI tools to multiply your own research velocity
  • Solid engineering skills in Python and deep learning frameworks (e.g., PyTorch), with the ability to work with large-scale video data pipelines
Nice to Have
  • Publications at top venues (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or similar)
  • Experience with video understanding, long-context temporal modeling, or efficient inference for edge/cloud deployment
  • Experience deploying ML models in consumer products at scale
Compensation

The base pay range for this role is $134,000 – $181,000 per year.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Vision-Language AI Scientist for Smart Home Video
Senior Vision-Language AI Scientist for Smart Home Video

Wyze • Kirkland (WA)

On-site
USD 134,000 - 181,000
Research Scientist - Vision Language Model
Research Scientist - Vision Language Model

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Member of Technical Staff, Vision / Language
Member of Technical Staff, Vision / Language

xdof.ai • San Mateo (CA)

On-site
USD 120,000 - 160,000
Competitive compensation and equity
Comprehensive health and wellness benefits
Collaborative and fast-paced work environment
Research Engineer, Visual Knowledge Work
Research Engineer, Visual Knowledge Work

Anthropic • New York (NY)

Hybrid
USD 350,000 - 850,000
Generous vacation
Parental leave
Flexible working hours
Member of Technical Staff, Vision/Language
Member of Technical Staff, Vision/Language

XDOF • San Mateo (CA)

On-site
USD 130,000 - 210,000
Research Engineer, Multimodal Data
Research Engineer, Multimodal Data

Eventual • San Francisco (CA)

On-site
USD 120,000 - 150,000
Catered lunches and dinners
Commuter benefit
Health, vision, and dental coverage
+2
Research Scientist, Video Understanding & World Models
Research Scientist, Video Understanding & World Models

Kindredventures • New York (NY)

On-site
USD 120,000 - 150,000
Research Vision Expertise
Research Vision Expertise

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer, Multimodal Data
Research Engineer, Multimodal Data

Eventual • United States

On-site
USD 150,000 - 210,000
In-person, tight-knit SF team
Competitive compensation
Startup equity
+7
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Clearview AI • United States

On-site
USD 180,000 - 250,000
Medical plans
Dental plans
Vision plans
+1