R&D Engineer / Postdoc — Vision-Language Models for Monitoring on Existing Cameras

Camly AI

Villeurbanne

On-site

EUR 36,000 - 48,000

Full time

7 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Subsidised meals
Transport reimbursement
Vocational training

Job summary

Camly AI, hosted within the Inria Startup Studio in Lyon, seeks a Postdoctoral researcher in R&D to advance vision-language models for monitoring with existing cameras. You will design evaluation protocols, adapt open-weight VLMs, and curate surveillance datasets to improve reliability.

You will work closely with the founder, publish internal reports, and contribute to product-ready insights while addressing privacy and European deployment constraints.

Qualifications

  • PhD in computer vision, machine learning or a related field.
  • Experience with multimodal or vision-language models and evaluation.
  • Experience building or curating datasets and using annotation tools.

Responsibilities

  • Define evaluation tasks and metrics for vision-language models on monitoring.
  • Fine-tune and evaluate open-weight models for Camly's surveillance use cases.
  • Curate and re-annotate public surveillance datasets relevant to logistics and safety.
  • Write internal technical reports and present results to the team and partners.
  • Collaborate with the founder and participate in pilot site visits when needed.

Skills

Multimodal models
Evaluation protocols
Fine-tuning
Dataset curation

Education

PhD in computer vision

Tools

CVAT
Label Studio

Job description

R&D Engineer / Postdoc — Vision-Language Models for Monitoring on Existing Cameras
  • Contract type: Fixed-term (10 months)
  • Required degree: PhD or equivalent
  • Position: Postdoctoral researcher
Context and assets of the position

This position is offered within the Inria Startup Studio programme at the Inria Centre at Lyon, as part of the Camly AI startup project.

The context:

Warehouses, industrial sites and buildings already have cameras. They are used almost only to review footage after an incident. Meanwhile, risky situations settle in and last: a blocked emergency exit, an obstructed aisle, safety equipment made inaccessible.

Existing video analytics fit this need poorly. They often rely on detectors trained case by case, which are costly to adapt, or on intrusive technologies (facial recognition, people tracking) that are hard for employees and regulators to accept.

Camly AI takes a different approach. Vision-language AI agents connect to existing cameras, with no new hardware. Users describe in plain language what to watch; the system flags abnormal situations that persist and produces reports useful for safety and compliance. Privacy is a design principle: no facial recognition, no biometrics, no individual tracking. The project is moving: the core pipeline works, the first market is chosen (logistics warehouses in the Lyon region. The project is led by its founder, a PhD in computer vision with a background in industrial automation. What is missing is the research strength that will make the system measurably reliable on real surveillance footage: could that be you?

Assignment

Vision-language model evaluation. Build a rigorous evaluation protocol for VLMs on monitoring tasks. Compare commercial and open-weight models on accuracy, false alarms, robustness to real conditions (lighting, high-angle views, cluttered scenes) and inference cost.

Adaptation and fine-tuning. Adapt open-weight VLMs to Camly's tasks (fine-tuning, parameter-efficient adaptation), with a view to sovereign deployment on European cloud or on-premise.

Datasets. Select, curate and re-annotate public surveillance datasets.

Transfer to the product. Turn results into concrete product improvements, working directly with the founder, and occasionally join pilot site visits.

Main activities
  • Define reference tasks and evaluation metrics (precision, recall, false-alarm rate, cost per camera).
  • Set up a reproducible evaluation bench and compare several VLM families on it.
  • Survey, select and re-annotate public datasets relevant to logistics and industrial safety.
  • Fine-tune and evaluate open-weight models, documenting gains and limits.
  • Write internal technical reports.
  • Present results to the team, to partners and, where relevant, to pilot sites.
Skills

Technical skills and level required:

  • PhD in computer vision, machine learning or a related field.
  • Proven experience with multimodal or vision-language models: evaluation, prompting, fine-tuning (for example Qwen-VL, InternVL, LLaVA, PaliGemma).
  • Experience building or curating datasets and using annotation tools (CVAT, Label Studio or equivalent).

Interpersonal skills:

  • Autonomy and a drive to build in a startup environment.
  • Ability to explain technical results to non-specialists.
  • Team spirit, rigour and honesty about the limits of results.
  • Subsidised meals
  • Partial reimbursement of public transport costs
  • Access to vocational training
Remuneration

According to the Inria salary scale ISS (inria startup studio).

General information
  • Desired start date: as soon as possible

Defence security: this position is likely to be located in a restricted access area (ZRR), as defined in Decree No. 2011-1425 on the protection of the nation's scientific and technical potential (PPST). Access is granted by the head of the institution after a favourable ministerial opinion; an unfavourable opinion would cancel the recruitment.

  • Recruitment policy: as part of its diversity policy, all Inria positions are open to people with disabilities.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer: Detailed riggable humans from multi-view video
Research Engineer: Detailed riggable humans from multi-view video

Inria, the French national research institute for the digital sciences • France

Hybrid
EUR 26,000 - 35,000
Transport partiellement remboursé
7 semaines de congés + RTT
Télétravail (90 jours/an)
+4
Post-Doctoral Research Visit F/M Crowd dynamics data acquisition and processing for large-scale dataset construction
Post-Doctoral Research Visit F/M Crowd dynamics data acquisition and processing for large-scale dataset construction

Inria • Rennes

On-site
EUR 26,000 - 36,000
Transport partiel remboursé
Congés annuels: 7 semaines + RTT
Télétravail après 6 mois
+3
CENTRALE LYON - Post-doctoral researcher (12 months) Mathematics of energy-efficient AI: diffusion models on emerging hardware
CENTRALE LYON - Post-doctoral researcher (12 months) Mathematics of energy-efficient AI: diffusion models on emerging hardware

Centrale Lyon • Écully

On-site
EUR 32,000 - 42,000
Chercheur·euse en deep learning / computer vision H/F
Chercheur·euse en deep learning / computer vision H/F

Adoc Talent Management • Paris

On-site
EUR 60,000 - 80,000
Rémunération attractive
BSPCE (options de stock)
Research Engineer: Detailed riggable humans from multi-view video
Research Engineer: Detailed riggable humans from multi-view video

Inria • France

Hybrid
EUR 2,600 - 3,600
Partial transport reimbursement
7 weeks leave + RTT days + exceptional
90 days teleworking per year
+4
Research Engineer F/M — Crowd data acquisition, processing and modelling
Research Engineer F/M — Crowd data acquisition, processing and modelling

Inria • Rennes

On-site
EUR 40,000 - 70,000
Partial transport reimbursement
Annual leave + RTT days
Teleworking after 6 months
+4
Chercheur·euse en deep learning / computer vision H/F
Chercheur·euse en deep learning / computer vision H/F

Adoc Tm • Paris

On-site
EUR 50,000 - 80,000
Rémunération attractive
BSPCE
Télétravail régulier
CENTRALE LYON - Post-doctoral researcher (12 months) Mathematics of energy-efficient AI: diffusion models on emerging hardware
CENTRALE LYON - Post-doctoral researcher (12 months) Mathematics of energy-efficient AI: diffusion models on emerging hardware

ecolecentraledelyon • Écully

On-site
EUR 24,000 - 36,000
Double supervision
Conference travel support
Housing assistance for internationalRe
Master Internship in AI for biological microscopy
Master Internship in AI for biological microscopy

Inria • France

Hybrid
EUR 12,000 - 17,000
Teleworking possible
Flexible working hours
Professional equipment provided
+1
Post-doctorant en mesure du réalisme des environnements 3D virtuels pour l'apprentissage - SimFi-Ed (F/H) / Postdoctoral Researcher in Assessing the Realism of Virtual 3D Environments for Learning - SimFi-Ed (W/M)
Post-doctorant en mesure du réalisme des environnements 3D virtuels pour l'apprentissage - SimFi-Ed (F/H) / Postdoctoral Researcher in Assessing the Realism of Virtual 3D Environments for Learning - SimFi-Ed (W/M)

Univ Lemans • Laval

On-site
EUR 34,000 - 38,000