AI Research Engineer — Vision & VLMs for Real-World AI

Palona AI

New York (NY)

On-site

USD 140,000 - 190,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive salary
Stock options
Green card sponsorship for qualified U
Medical, dental, vision benefits

Job summary

Palona AI is building AI for the physical world, starting with restaurants. We seek an AI Research Engineer to advance visual intelligence for image/video understanding and multimodal reasoning in real restaurant environments.

You will own research questions, build datasets, train and evaluate models, and partner with product and engineering to deploy successful approaches into production. We encourage researchers from autonomous driving, robotics, and embodied AI to apply.

Qualifications

  • 3+ years of research or applied development in computer vision or multimodal learning (or relevant grad research).
  • Strong track record in CV or vision-language modeling via publications, projects, or industry work.
  • Strong foundations in deep learning, visual representation learning, and experimental design.
  • Hands-on experience training or adapting CV models and evaluating VLMs beyond API usage.
  • Strong Python skills and PyTorch or equivalent DL frameworks.
  • Experience building datasets, designing evaluations, and using ablations to understand improvements.
  • Ability to turn research code into reproducible, tested systems for others to use.
  • Ability to align modeling choices with product constraints like latency and privacy.

Responsibilities

  • Develop CV and VLM approaches for scene understanding, object detection and tracking, activity recognition, and video event understanding.
  • Adapt, fine-tune, and evaluate vision and vision-language models for grounding, temporal reasoning, and structured prediction.
  • Design training strategies including supervised fine-tuning, representation learning, distillation, and domain adaptation.
  • Build representative image/video datasets, annotation workflows, and evaluation sets capturing edge cases while protecting data.
  • Create rigorous experiments and benchmarks measuring perception quality, robustness, latency, and costs across locations.
  • Diagnose failures from occlusion, lighting, camera placement, rare events, and domain shift; improve data and models.
  • Collaborate with infra and product engineers to deploy efficient inference pipelines with monitoring and rollback paths.
  • Translate advances into practical product capabilities and communicate tradeoffs and evidence behind decisions.
  • Raise standards through reproducible experiments, reviews, and documentation.

Skills

Python
Deep learning
Computer vision
Multimodal learning
Experiment design
Research

Education

PhD or related master in CV / ML / robotics

Tools

PyTorch

Job description

Palona AI is building AI for the physical world, starting with restaurants. We seek an AI Research Engineer to advance visual intelligence for image/video understanding and multimodal reasoning in real restaurant environments.

You will own research questions, build datasets, train and evaluate models, and partner with product and engineering to deploy successful approaches into production. We encourage researchers from autonomous driving, robotics, and embodied AI to apply.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Engineer: Vision & VLMs (Stock Options)
AI Research Engineer: Vision & VLMs (Stock Options)

Palona AI • Los Altos (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Stock options
Green card sponsorship
+3
AI Research Engineer, Computer Vision & VLMs
AI Research Engineer, Computer Vision & VLMs

Palona AI • Los Altos (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Stock options
Green card sponsorship
+3
AI Research Engineer, Computer Vision & VLMs
AI Research Engineer, Computer Vision & VLMs

Palona AI • New York (NY)

On-site
USD 140,000 - 190,000
Competitive salary
Stock options
Green card sponsorship for qualified U
+1
Research Engineer
Research Engineer

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 140,000 - 190,000
Vision AI & 3D ML Research Engineer (Remote)
Vision AI & 3D ML Research Engineer (Remote)

Centific • Seattle (WA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Product Engineer - End-to-End Restaurant Tech
AI Product Engineer - End-to-End Restaurant Tech

Palona AI • Los Altos (CA)

On-site
USD 140,000 - 190,000
Competitive salary and stock options
Health, dental, vision benefits
Family leave
+3
Research Engineer - Vision Language Models / Multimodal AI / Computer Vision
Research Engineer - Vision Language Models / Multimodal AI / Computer Vision

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 180,000 - 240,000
AI Robotics Engineer, Vision-Language-Action (VLA)
AI Robotics Engineer, Vision-Language-Action (VLA)

Confidential • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI Infra Engineer - Real-Time Vision & LLMs
AI Infra Engineer - Real-Time Vision & LLMs

Ambient • Redwood City (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Stock options
Health + welfare package
Flexible time off
+1
Lead Computer Vision Engineer for Real-Time Retail AI
Lead Computer Vision Engineer for Real-Time Retail AI

United States Digital Space LLC • United States

Hybrid
USD 180,000 - 260,000
Flexible work locations