Research Intern, Small Language Models (Mitacs)

Onix (doing business as “Onix”; legal entity listed as 16445039 Canada Inc.)

Montreal (administrative region)

Hybrid

CAD 17,000 - 23,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Onix is seeking a graduate student or postdoc for a Mitacs Accelerate research internship in Montreal. You will train per-expert SLMs, work across the training stack, and refine fidelity evals while building the data layer feeding private corpora and grounded retrieval.

You will collaborate closely with our R&D and AI team, CTO, and academic supervisor. The role emphasizes on-premise work from Old Montreal with strict privacy boundaries.

Qualifications

  • Current grad student or postdoc at a Canadian institution.
  • Extensive work in small language models, efficient training, distillation, retrieval, grounding, or on-device inference.
  • Experience shipping or publishing in the relevant AI/ML space.

Responsibilities

  • Train and fine-tune per-expert SLMs with distillation, RLHF/RLVR, and Expert Fidelity evals.
  • Collaborate across the SLM training stack: corpus, training, evals, deployment.
  • Improve Expert Fidelity evals and push models to clear the bar.
  • Develop data layer components: corpus ingestion, chunking, agentic RAG, grounding.
  • Ship to production during the internship with real users and real corpora.

Skills

SLM training
RLHF/RLVR
Per-expert isolation
Data-to-eval path
Production readiness

Education

Graduate student / Postdoc

Tools

PyTorch
HuggingFace
RAG
On-device inference

Job description

You will work on the end-to-end pipeline that turns one expert’s private corpus into a small language model that is faithful to them, not just in voice but in what it knows, what it cites, and the judgment calls it makes. The expert grades it.

Onix is the Personal Intelligence platform. Each onix is a small language model trained exclusively on a single expert’s private corpus: clinical notes, unpublished research, proprietary methods that have never been on the internet and never will be. Fully isolated, never bleeding across experts.

Privacy is an engineering surface, not a compliance line. Per-expert isolation, on-device inference paths, federated update strategies, and grounding guarantees are open research areas you will work on.

Most of the work is SLM training: per-expert fine-tuning, distillation, preference learning, and the Expert Fidelity evals that gate every release, scoring each model on accuracy, groundedness, persona, and judgment. You will also work on the data layer that feeds it: ingesting and chunking a private corpus, and the agentic RAG that grounds the model. Training is the core. The data and retrieval layer is where that training meets the product.

Every session generates refinement signal. Every expert validates outputs in the loop. This is exclusive, expert-graded data that Big AI cannot scrape, replicate, or buy.

This is a Mitacs Accelerate research internship. It is project-scoped, four or six months to start, with the option to extend. You work in person from our office in Old Montreal, collaborating closely with our R&D and AI team, our CTO, and your academic supervisor. We are recruiting primarily from Mila, and open to any grad student or postdoc eligible for Mitacs.

The next great AI lab will prove that human expertise is a moat, not a training set. That is what we are proving.

What You Will Do
  • Train and fine-tune per-expert SLMs alongside the R&D team: distillation, preference learning (RLHF / RLVR), and the Expert Fidelity eval loop that gates them. This is the bulk of the internship.
  • Work across the SLM training stack with the team: corpus, training, evals, deploy.
  • Help sharpen the Expert Fidelity evals, and push the model to clear the bar.
  • Work on the data layer that feeds the model: ingest and chunk a private corpus, and the agentic RAG and grounding around it. Training is the core; this data layer is where your work touches the product.
  • Ship to production during the internship. Real experts. Real corpora. Real users. Real consequences if we get it wrong.
What You Will Work With
  • Training stack: modern PyTorch and HuggingFace tooling for accelerated per-expert fine-tuning. Distillation pipelines that generate training data without exposing a private corpus. Expert Fidelity evals as the gating bar.
  • Preference learning at the per-expert level: RLHF where each expert is the literal human in the loop, and RLVR grounded in our Expert Fidelity reward surface. The data and the experts are exclusive.
  • Inference stack: per-expert endpoints. You work on quantization, caching, and the per-expert latency budget.
  • On-device and edge deployment for iOS. Latency and battery are first-class constraints, not afterthoughts.
  • The data layer that feeds training: corpus ingestion, document chunking, agentic RAG, and grounding over a single expert’s corpus. Mostly SLM training, but your work here has direct product impact.
  • Per-expert isolation infrastructure on real users from day one. No data crosses expert boundaries. Privacy is architecture.
Who You Are

You can read a paper, prototype the model, and ship it to production in the same week. You are a current grad student or postdoc, primarily from Mila, with substantive work in small language models, efficient training, distillation, retrieval, RAG, grounding, or on-device inference. You want your next system to ship to real users instead of a benchmark.

You came here for the mission. Technology should amplify human genius, not replace it.

We publish on our timeline, not a journal’s.

You have:

  • ML technical chops across training, eval, and retrieval.
  • Substantive shipped or published work in SLMs, distillation, RAG, retrieval, grounding, or on-device inference.
  • A track record of defending technical positions on open research questions with evidence, not vibes.
  • Comfort working across the whole data-to-eval path, not just the model.
  • Eligibility for a Mitacs internship: registered as a grad student or postdoc at a Canadian institution.

You are:

  • A research engineer first. Engineers here do research. Researchers here do engineering.
  • Opinionated. You can articulate a position on three open research questions in our space within an hour of joining.
  • Direct. You tell teammates they are wrong when they are, and accept the same in return.
  • In Old Montreal in person for the internship.

We do not care which lab you came from. We care what you have shipped, what you have measured, and what you would publish next.

What Success Looks Like
  • A per-expert SLM you helped train is in production and passes an Expert Fidelity bar an expert signs off on.
  • You moved a Fidelity eval metric the team cares about.
  • The data and retrieval layer you worked on measurably improves grounding for a live expert.
  • You publish or present what you built with the team.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Small Language Models
Member of Technical Staff, Small Language Models

Purchaseyourfreedom • Montreal (administrative region)

On-site
CAD 100,000 - 130,000
Machine Learning Intern/Co-op (Winter 2027)
Machine Learning Intern/Co-op (Winter 2027)

Cohere • Toronto

On-site
GBP 25,000 - 35,000
Weekly lunch stipend
Health and dental benefits
100% parental leave top-up
+2
ML Engineering Intern - Infrastructure (Spring 2027)
ML Engineering Intern - Infrastructure (Spring 2027)

Jaide Health • Toronto

On-site
CAD 28,000 - 39,000
Flexible PTO
ML Engineering Intern - Motion Capture (Spring 2027)
ML Engineering Intern - Motion Capture (Spring 2027)

Jaide Health • Toronto

On-site
CAD 21,000 - 32,000
Mentorship from world-class engineers
Hands-on production exposure
Flexible PTO
ML Engineering Intern - Reconstruction (Spring 2027)
ML Engineering Intern - Reconstruction (Spring 2027)

Jaide Health • Toronto

On-site
CAD 17,000 - 23,000
Flexible PTO
ML Engineering Intern - Reconstruction (Spring 2027) Onsite (Toronto, Canada)
ML Engineering Intern - Reconstruction (Spring 2027) Onsite (Toronto, Canada)

S27a • Toronto

On-site
CAD 32,000 - 42,000
AI Product Manager
AI Product Manager

INTO Inc. • Montreal (administrative region)

On-site
CAD 166,000 - 250,000
Remote work
Bonus program
Internal training
+1
ML Engineering Intern - Infrastructure (Spring 2027) Onsite (Toronto, Canada)
ML Engineering Intern - Infrastructure (Spring 2027) Onsite (Toronto, Canada)

S27a • Toronto

On-site
CAD 20,000 - 31,000
Flexible PTO
Senior Software Engineer
Senior Software Engineer

Polarized Capital • Montreal (administrative region)

On-site
CAD 150,000 - 200,000
Machine Learning Engineering Intern - Motion Capture (Fall 2026)
Machine Learning Engineering Intern - Motion Capture (Fall 2026)

S27a • Toronto

On-site
CAD 25,000 - 35,000
Flexible Paid Time Off