Location: Brussels, hybrid · Type: Full-time · Reports to: CTO · Team: you lead the AI track (growing under you)
The thesis
Lead AI Scientist — Computer Vision
Location: Brussels, hybrid · Type: Full-time · Reports to: CTO · Team: you lead the AI track (growing under you)
iFreq is building football's measurement standard: every number protocol-certified, video-proven, and comparable across time, age and geography. Two layers make it work - a certified score that institutions demand (proctored, protocol-bound, portable) and a practice tier measured by a phone camera every day. A hard schema wall keeps them apart: practice predicts, certification proves.
Both layers stand on one capability: a camera that can score a football drill. Under protocol conditions it certifies a flagship drill; in loose conditions it scores a bedroom drill. Our roadmap calls phone-camera CV \"not a cost-reduction feature but the platform keystone ... both halves of the flywheel depend on it - resource it first.\" This role is that resourcing.
You will author and own the computer-vision R&D programme - the research plan, its milestones, its error budgets - and then run it, from the first labelled rep to models that run offline on a mid-range Android.
Where we are (September 2026)
- A first model is in production. In August 2026 we shipped the first iFreq AI video model to our analyst review platform: temporal action localisation that segments on-ball / off-ball phases and pre-fills annotations for the analysts who collect raw data from drill videos. It covers three drills across seven level variants (BTL L1, DYZ L1-L3, UPL L1-L3).
- A validation layer is being built on top of it - automatic flags for implausible annotations, because analysts get tired and corrupted data must be caught before it reaches a score.
- The platform around the model exists. A video_analysis domain in the backend with a processing queue, stored segment timelines, and automatic metrics kept in their own table so manual and automatic values stay comparable; an on-demand GPU host that wakes when footage is queued and stops itself when the queue runs dry; a pluggable run_inference() seam that is literally waiting for the model code.
- The data is real and growing. Tens of thousands of rep clips, each bound to a player, an exercise, a level and a rep index; per-rep human-entered metrics; and a review timeline that merges analyst and model annotations by authority so unreviewed spans are never silently labelled \"nothing\".
- Scoring is versioned. Every score carries scoring_version and scoring_uid provenance; the coach pentagon is a player-scoped rolling model with fatigue-weighted reps and hybrid z-benchmarks. Your models will feed it, not replace it.
What you will do
Own the research programme
- Write and maintain the CV R&D plan: problem decomposition, milestones, error budgets per drill, build / buy / defer decisions, and the evidence that moves a date.
- Decide what gets modelled and how: temporal action localisation, pose estimation, ball and player tracking, per-rep event detection - chosen against what the protocol actually needs to measure.
- Represent the science externally: reliability and validity publications, the sport-science advisory board, and the federations and academies that will ask \"how do you know?\"
Ship the models (iFreq 2)
- Take the assisted-review model from suggestion to automatic rep detection across all eight exercises of the catalog.
- Build phone-camera timing for the flagship drills and prove it against timing gates under protocol conditions, with an error budget you set and defend.
- Move models to the edge: quantised, on-device inference for the daily session and the free baseline test, with the latency and battery budgets a mid-range Android imposes.
Run measurement science, not just ML
- Design labelling protocols, inter-rater agreement checks and dataset versioning so ground truth is a maintained asset, not a folder.
- Compute and publish test–retest reliability (ICC) per drill, calibration and uncertainty on every automatic value - \"82 ± 3, from 4 certified sessions\" is the product, not a footnote.
- Build the anomaly-defense layer: impossible values, duplicate reps, operator drift, and later cross-angle disagreement - flagged, never silently corrected.
Productionise with the platform team
- Fill the run_inference() seam on the GPU worker and own the model lifecycle behind it: weights, versions, regression tracking across scoring_versions.
- Work with the backend lead on the practice/certified schema wall, with the mobile lead on on-device capture and inference, and with product on what a coach or parent sees when a model is unsure.
Lead the AI track
- Grow the team from one AI engineer to the group the programme needs; set the evaluation bar; review the work.
- Make the work legible to a five-person tech team and a founder-led company: short written updates, decisions with reasons, demos over decks.
You should have
- A research track record in computer vision with depth in video understanding: temporal action localisation or detection, pose estimation, tracking.
- You know the distance between a benchmark number and a deployed one.
- Measurement-science instincts: reliability, agreement, calibration and uncertainty are familiar tools (ICC, Bland Altman, conformal or Bayesian intervals), and you are comfortable publishing them.
- Edge deployment experience: quantisation, ONNX / TFLite / Core ML, and the discipline of a latency budget on a mid-range phone.
- Pragmatic data engineering: labelling workflows, active learning, dataset versioning, and the patience to build the ground truth before the model.
- Deep PyTorch and Python; comfort shipping behind a real API (we run FastAPI, PostgreSQL and AWS) rather than in notebooks.
- Leadership of a small research or applied team, and writing clear enough that an R&D programme can be read and funded on the strength of the document.
- Professional English. French is a strong plus — the team and most of our coaches work in it.
Nice to have
- Sport science or biomechanics: movement screening, athletic testing protocols, youth maturation (PHV, relative-age effects).
- Self-supervised or foundation video models applied to small, domain-specific datasets.
- Experience writing and executing funded R&D programmes (regional or national innovation grants).
- (iFreq 3) Experiment tracking, model registries and drift monitoring in production.
What the first twelve months look like
- By end of 2026: the validation layer is live on the review platform; ground truth is versioned with published inter-rater agreement; auto rep-detection is running as a prototype on the GPU worker.
- Q1 2027: ICC per drill is computed from our own data and published; auto rep-detection measurably reduces analyst time per video on the covered drills.
- Q2-Q3 2027: phone-camera timing for the first flagship drill passes its error budget against timing gates under protocol conditions; the practice-tier scoring proof-of-concept runs in loose conditions.
- By Sep 2027: the R&D plan for on-device inference and the daily session is written, resourced and started.
How we work
Small team, real data, direct line to the founders. Tech is five people plus the CTO; AI is a first-class track with its own ownership and budget, not a service function. Our science is published rather than claimed, our scores carry provenance, and a wrong number shown to one parent is treated as a trust bug, not tech debt.