AI Research Engineer, Computer Vision & VLMs

Palona AI

Toronto

On-site

CAD 120,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary and stock option
Green card sponsorship for eligible U
Medical, dental, vision benefits
Family leave and disability benefits
Paid time off and company holidays
Learning and development support

Job summary

Palona AI in Toronto is seeking an AI Research Engineer with a strong background in computer vision and vision-language models (VLMs) to develop the visual intelligence behind Palona’s products. You will work on image and video understanding, spatiotemporal reasoning, and multimodal models for actionable insights in real restaurant environments.

You will own end-to-end research-to-production work, from dataset creation and model training to partnering with product and engineering to deploy

Qualifications

  • 3+ years of research or applied development in computer vision or multimodal learning.
  • Record of publications or impactful research projects in CV/VLMs.
  • Strong foundations in deep learning, representation learning, and experimental design.
  • Experience training or adapting CV models and evaluating VLMs beyond API use.
  • Excellent Python and modern DL tooling skills.
  • Experience building datasets and reliable evaluations with ablations.

Responsibilities

  • Develop CV and VLM approaches for scene understanding, detection, tracking, and event recognition in video.
  • Fine-tune and evaluate vision and vision-language models for grounding and temporal reasoning.
  • Design training strategies, including fine-tuning, representation learning, distillation, and adaptation.
  • Build datasets, annotation workflows, and evaluation sets including edge cases while protecting data.
  • Run rigorous experiments and benchmarks on perception quality, latency, and robustness.
  • Diagnose failures due to occlusion, lighting, or domain shift and improve data/models.
  • Collaborate with infra and product teams to deploy efficient inference pipelines.
  • Communicate tradeoffs and evidence behind modeling decisions to stakeholders.
  • Raise research/engineering standards with reproducible experiments and documentation.

Skills

Computer vision
Vision-language models
Python
PyTorch
Deep learning
Experiment design
Dataset construction

Education

PhD or research-focused Master’s in CV/ML

Tools

PyTorch

Job description

Palona is building AI for the physical world, starting with restaurants. Understanding a busy restaurant means making sense of people, objects, activities, and events as they change over time, despite occlusion, changing lighting, varied camera views, and incomplete information.

We are looking for an AI Research Engineer with a strong research background in computer vision and vision-language models (VLMs) to develop the visual intelligence behind Palona’s products. You will work on image and video understanding, spatiotemporal reasoning, and multimodal models that connect visual observations to useful insights and actions in real restaurant environments.

This role combines research depth with ownership of working systems. You will formulate research questions, build datasets, train and evaluate models, and partner with product and engineering to bring successful approaches into production. Researchers and engineers from autonomous driving, robotics, embodied AI, and related perception fields are especially encouraged to apply.

What you’ll own
  • Develop computer vision and VLM approaches for scene understanding, object detection and tracking, activity recognition, and understanding events across video.
  • Adapt, fine-tune, and evaluate vision and vision-language models for visual grounding, temporal reasoning, and structured prediction grounded in observable evidence.
  • Design training and adaptation strategies, including supervised fine-tuning, representation learning, distillation, and domain adaptation, based on measurable product needs.
  • Build representative image and video datasets, annotation workflows, and evaluation sets that capture difficult edge cases while protecting sensitive data.
  • Create rigorous experiments and benchmarks that measure perception quality, temporal consistency, hallucinations, robustness, latency, and cost across locations and operating conditions.
  • Diagnose failures caused by occlusion, lighting changes, camera placement, rare events, and domain shift; use those findings to improve data and models.
  • Partner with infrastructure and product engineers to deploy efficient inference pipelines, with monitoring, quality gates, staged rollouts, and rollback paths.
  • Translate advances in computer vision, VLMs, and embodied AI into practical product capabilities, and communicate the evidence and tradeoffs behind your decisions.
  • Raise research and engineering standards through reproducible experiments, thoughtful reviews, and clear documentation.
  • 3+ years of research or applied development experience in computer vision, multimodal learning, or a closely related field; relevant graduate research counts toward this experience.
  • A demonstrated research track record in computer vision or vision-language modeling, through publications, substantial research projects, open-source contributions, or research delivered in industry.
  • Strong foundations in deep learning, visual representation learning, and experimental design, with depth in areas such as video understanding, detection and tracking, visual grounding, or multimodal reasoning.
  • Hands-on experience training, fine-tuning, or adapting computer vision models, and developing or evaluating VLMs beyond basic API integration.
  • Strong Python skills and experience with PyTorch or an equivalent deep learning framework, along with modern training and evaluation tooling.
  • Experience building datasets, designing reliable evaluations, analyzing model failures, and using ablations to understand what drives improvements.
  • Strong software engineering judgment and the ability to turn research code into reproducible, tested systems that other engineers can use.
  • Ability to connect modeling choices to product constraints including latency, cost, privacy, reliability, and user experience.
  • Comfort working through ambiguity and collaborating across research, engineering, and product.
Especially relevant experience
  • A PhD or research-focused master’s degree in computer vision, machine learning, robotics, or a related field, or equivalent research experience.
  • Industry research or engineering experience in autonomous driving, robotics, embodied AI, or other applications of perception in the physical world.
  • Publications at venues such as CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, CoRL, ICRA, or RSS.
  • Experience with monocular video perception, spatial understanding, long-video reasoning, or learning from limited and noisy labels.
  • Experience shipping vision models under real-time constraints, including model compression, distillation, quantization, or inference optimization.
  • Competitive salary and stock option plan.
  • Company-sponsored green card applications for strong candidates hired into U.S.-based roles, subject to eligibility.
  • Medical, dental, vision, and retirement benefits as applicable.
  • Family leave and short-term and long-term disability benefits as applicable.
  • Paid time off and company holidays.
  • Learning and development support.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Engineer, Computer Vision & VLMs
AI Research Engineer, Computer Vision & VLMs

Palona • Toronto

On-site
CAD 140,000 - 190,000
Competitive salary
Stock options
Green card sponsorship
+4
AI Modeling Engineer
AI Modeling Engineer

Palona AI • Toronto

On-site
CAD 120,000 - 160,000
Stock options
Benefits package
Family leave
+3
AI Infrastructure Engineer
AI Infrastructure Engineer

Palona AI • Toronto

On-site
CAD 90,000 - 130,000
Competitive Salary
Stock Option Plan
Medical benefits
+4
AI Software Engineer, Growth
AI Software Engineer, Growth

Palona AI • Toronto

On-site
CAD 100,000 - 140,000
Stock options
Health benefits
Family leave
+2
MLOps & Data Engineer
MLOps & Data Engineer

Palitronica Inc. • Southwestern Ontario

On-site
CAD 100,000 - 150,000
Member of Technical Staff, Research Engineer
Member of Technical Staff, Research Engineer

Moonvalley • Toronto

On-site
CAD 110,000 - 170,000
Competitive salary and equity
Private health coverage
Pension contribution
+4
Machine Learning Engineer
Machine Learning Engineer

Invision AI Inc. • Toronto

On-site
CAD 110,000 - 160,000
Equity
Sr. Computer Vision Engineer
Sr. Computer Vision Engineer

OpenSpace • Toronto

On-site
CAD 168,000 - 220,000
Base salary + benefits
Career development
Growth opportunities
Senior Data Scientist
Senior Data Scientist

Alberta Veterinary Laboratories Ltd • Southwestern Ontario

On-site
CAD 110,000 - 170,000
Senior Software Engineer, Adaptive
Senior Software Engineer, Adaptive

Miovision • Canada

Hybrid
CAD 110,000 - 170,000
Health benefits
RRSP matching
Flexible vacation
+3