Machine Learning Data Engineer

kadence

Montreal (administrative region)

On-site

CAD 120,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Kadence in Montreal is hiring multiple Senior Machine Learning Data Engineers to join an ambitious AI research team focused on training-data engineering and curation at web-scale. You’ll help transform raw web-scale data into high-quality datasets used to train next-generation models.

This role sits at the intersection of data engineering, ML, NLP, and large-scale training data curation, building scalable infrastructure, quality scoring, and tooling for researchers.

Qualifications

  • 5+ years of experience across ML engineering, data processing, NLP, or data infrastructure.
  • Experience with very large unstructured text datasets, ideally web/foundation-model scale.
  • Strong production-level Python programming.
  • Experience with distributed data-processing technologies such as Spark, Ray, Flink.
  • Experience with PII scrubbing, content-safety filtering, dataset quality.

Responsibilities

  • Design, build and scale pipelines that transform raw web-scale data into high-quality datasets for training large ML models.
  • Process extremely large unstructured text corpora, including datasets at trillion-token scale.
  • Build data-processing systems covering deduplication, model-based quality scoring, filtering, PII removal, metadata extraction.
  • Develop automated data-quality systems using heuristics, ML classifiers, LLM evaluators and human-in-the-loop workflows.
  • Build monitoring and guardrails to identify data-quality regressions before they impact training.
  • Collaborate with AI researchers to understand evolving data requirements and identify gaps in datasets.
  • Design robust contamination and data-leakage detection systems.
  • Build internal tools to explore, query and understand large datasets.
  • Optimize large-scale distributed processing for throughput, reliability and cost.

Skills

Python
Distributed data processing
Spark
Ray
Flink
Data pipelines
NLP
Machine learning engineering
PII scrubbing
Airflow

Tools

Spark
Ray
Flink
Airflow
Prefect
Dagster
Docker
Kubernetes
Terraform / IaC
Experiment tracking

Job description

This is an opportunity to work alongside a highly accomplished AI research team tackling fundamental problems in advanced machine learning.

Rather than maintaining established pipelines, you’ll be helping develop new approaches to training-data engineering and curation where established playbooks often do not yet exist.

The organization is building a substantial technical team in Montreal, and we are hiring multiple people across this area.

If you’ve worked on LLM training data, large-scale NLP pipelines, foundation-model infrastructure or web-scale data processing, I’d be interested in speaking with you.

We’re working with an ambitious AI research organization building next-generation machine learning systems and are looking for multiple Senior Machine Learning Data Engineers to join its growing technical team.

This role sits at the intersection of data engineering, machine learning, NLP, and large-scale training data curation.

You’ll be responsible for building and scaling the infrastructure that transforms raw, web-scale data into high-quality datasets used to train advanced AI models.

The core challenge is engineering data quality at enormous scale: developing filtering systems, model-based quality scoring, dataset transformations, contamination detection, and tooling that allows researchers to understand and work with training corpora effectively.

What You’ll Work On
  • Design, build and scale pipelines that transform raw web-scale data into high-quality datasets for training large machine learning models.
  • Process extremely large unstructured text corpora, including datasets at trillion-token scale.
  • Build sophisticated data-processing systems covering: Deduplication, Model-based quality scoring, Heuristic filtering, Toxicity and content-safety filtering, PII detection and removal, Metadata extraction, Dataset transformations, Versioning and provenance tracking
  • Develop automated data-quality systems using a combination of heuristics, ML classifiers, LLM-based evaluators and human-in-the-loop workflows.
  • Build monitoring and guardrails to identify data-quality regressions before they impact downstream model training.
  • Work closely with AI researchers to understand evolving training-data requirements and identify gaps within existing datasets.
  • Design robust evaluation contamination and data-leakage detection systems.
  • Build internal tools that allow researchers to easily explore, query and understand large datasets.
  • Optimize large-scale distributed processing systems for throughput, reliability and infrastructure cost.
What We’re Looking For
  • 5+ years of experience across machine learning engineering, data processing, NLP, data infrastructure or related areas.
  • Experience working with very large unstructured text datasets, ideally at web or foundation-model scale.
  • Strong production-level Python.
  • Hands‑on experience with distributed data-processing technologies such as: Spark, Ray, Flink
  • Experience designing and optimizing high-throughput distributed pipelines.
  • Experience with areas such as: PII scrubbing, Content-safety filtering, Dataset quality, Evaluation contamination, Training-data governance
  • Experience with workflow orchestration frameworks such as Airflow, Prefect or Dagster.
  • Ability to work closely with both research and engineering teams and translate research requirements into scalable infrastructure.
Particularly Relevant Experience

We’d be especially interested in people who have worked on:

  • Pretraining or post-training data pipelines for LLMs
  • Common Crawl or comparable large-scale datasets
  • LLM-based data-quality evaluation
  • ML classifiers for filtering or scoring training data
  • Dataset deduplication and contamination detection
  • Distributed ML/data infrastructure

Experience with vLLM, SGLang, Docker, Kubernetes, infrastructure-as-code, experiment tracking or open-source NLP/data tooling would also be valuable.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer
Machine Learning Engineer

Altis Technology • Montreal (administrative region)

Hybrid
CAD 90,000 - 120,000
Exposure to complex, enterprise-scale machine learning initiatives
Opportunities with modern ML frameworks and cloud technologies
Collaborative environment that values innovation
+1
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Encore Technical Solutions Inc. • Toronto

On-site
CAD 100,000 - 150,000
ML Engineer – Generative AI & LLMs (Remote)
ML Engineer – Generative AI & LLMs (Remote)

Ample Insight Inc • Toronto

Hybrid
CAD 110,000 - 170,000
Data Engineer
Data Engineer

CoFoMo Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 140,000
Senior ML Data Processing Developer
Senior ML Data Processing Developer

LawZero • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Health benefits
20 days vacation
4% retirement contribution by employer
+2
Full Stack Engineer
Full Stack Engineer

Randstad Enterprise • Montreal (administrative region)

On-site
CAD 90,000 - 140,000
Machine Learning Engineer
Machine Learning Engineer

Linkus Group • Toronto

On-site
CAD 85,000 - 110,000
Senior Data Engineer - Machine Learning & Data Platforms - REMOTE
Senior Data Engineer - Machine Learning & Data Platforms - REMOTE

Talent To Hire Inc. • Toronto

Remote
CAD 100,000 - 130,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Motion Recruitment • Toronto

On-site
CAD 120,000 - 180,000
Senior AI Machine Learning Engineer
Senior AI Machine Learning Engineer

SAP SE • Montreal (administrative region)

On-site
CAD 108,000 - 223,000