AI Research Engineer: Vision & VLMs (Stock Options)

Palona AI

Los Altos (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive salary
Stock options
Green card sponsorship
Medical insurance
Paid time off
Learning & development

Job summary

Palona AI in the United States seeks an AI Research Engineer with a strong CV and vision-language background to build visual intelligence for restaurant environments. You will develop models for scene understanding, detection, tracking, and multimodal reasoning, and collaborate with product and engineering to deploy production-ready systems.

Ideal candidates have 3+ years of CV or multimodal research, a PhD or research-focused masters, and hands-on PyTorch experience.

Qualifications

  • 3+ years of research or applied development in computer vision or multimodal learning.
  • Demonstrated research track record in computer vision or vision-language modeling (publications, projects, or industry deliverables).
  • Strong foundations in deep learning, representation learning, and experimental design.
  • Hands-on experience training/fine-tuning computer vision models and evaluating VLMs.

Responsibilities

  • Develop CV and VLM approaches for scene understanding, object detection/tracking, activity recognition, and event understanding across video.
  • Adapt and evaluate vision-language models for grounding, temporal reasoning, and structured predictions grounded in observable evidence.
  • Design training and adaptation strategies including supervised fine-tuning, representation learning, distillation, and domain adaptation.
  • Build image/video datasets, annotation workflows, and evaluation sets capturing edge cases while protecting data privacy.
  • Create rigorous experiments/benchmarks measuring perception quality, robustness, latency, and cost across locations and conditions.
  • Collaborate with infrastructure and product engineers to deploy efficient inference pipelines with monitoring and rollback paths.

Skills

Computer vision
Vision-language modeling
Python
PyTorch
Deep learning
Dataset building
Model evaluation
Software engineering
Cross-functional collaboration
Edge-case analysis

Education

PhD in Computer Vision / ML
Master's degree in Computer Vision / ML

Tools

OpenCV
CUDA
Git
TensorBoard

Job description

Palona AI in the United States seeks an AI Research Engineer with a strong CV and vision-language background to build visual intelligence for restaurant environments. You will develop models for scene understanding, detection, tracking, and multimodal reasoning, and collaborate with product and engineering to deploy production-ready systems.

Ideal candidates have 3+ years of CV or multimodal research, a PhD or research-focused masters, and hands-on PyTorch experience.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Engineer — Vision & VLMs for Real-World AI
AI Research Engineer — Vision & VLMs for Real-World AI

Palona AI • New York (NY)

On-site
USD 140,000 - 190,000
Competitive salary
Stock options
Green card sponsorship for qualified U
+1
AI Research Engineer, Computer Vision & VLMs
AI Research Engineer, Computer Vision & VLMs

Palona AI • Los Altos (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Stock options
Green card sponsorship
+3
AI Research Engineer, Computer Vision & VLMs
AI Research Engineer, Computer Vision & VLMs

Palona AI • New York (NY)

On-site
USD 140,000 - 190,000
Competitive salary
Stock options
Green card sponsorship for qualified U
+1
Senior ML Engineer - Vision & Multimodal (Production)
Senior ML Engineer - Vision & Multimodal (Production)

Clearview AI • United States

On-site
USD 150,000 - 200,000
Medical, Dental, Vision
STD and LTD Plans
AI Robotics Engineer, Vision-Language-Action (VLA)
AI Robotics Engineer, Vision-Language-Action (VLA)

Confidential • San Francisco (CA)

On-site
USD 150,000 - 230,000
Research Engineer
Research Engineer

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 140,000 - 190,000
Research Engineer - Vision Language Models / Multimodal AI / Computer Vision
Research Engineer - Vision Language Models / Multimodal AI / Computer Vision

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 180,000 - 240,000
Research Scientist - VLM
Research Scientist - VLM

Storm3 • San Francisco (CA)

On-site
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+1
AI Modeling Engineer: Voice & Multimodal (Stock Options)
AI Modeling Engineer: Voice & Multimodal (Stock Options)

Palona AI • Los Altos (CA)

On-site
USD 140,000 - 210,000
Stock options
Benefits: medical/dental/vision/ret/le
Family leave
+3
Vision-Language Models (VLMs)
Vision-Language Models (VLMs)

TalentOla • Waukesha (WI)

On-site
USD 120,000 - 150,000