Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Kadence in Montreal is hiring multiple Senior Machine Learning Data Engineers to join an ambitious AI research team focused on training-data engineering and curation at web-scale. You’ll help transform raw web-scale data into high-quality datasets used to train next-generation models.
This role sits at the intersection of data engineering, ML, NLP, and large-scale training data curation, building scalable infrastructure, quality scoring, and tooling for researchers.
This is an opportunity to work alongside a highly accomplished AI research team tackling fundamental problems in advanced machine learning.
Rather than maintaining established pipelines, you’ll be helping develop new approaches to training-data engineering and curation where established playbooks often do not yet exist.
The organization is building a substantial technical team in Montreal, and we are hiring multiple people across this area.
If you’ve worked on LLM training data, large-scale NLP pipelines, foundation-model infrastructure or web-scale data processing, I’d be interested in speaking with you.
We’re working with an ambitious AI research organization building next-generation machine learning systems and are looking for multiple Senior Machine Learning Data Engineers to join its growing technical team.
This role sits at the intersection of data engineering, machine learning, NLP, and large-scale training data curation.
You’ll be responsible for building and scaling the infrastructure that transforms raw, web-scale data into high-quality datasets used to train advanced AI models.
The core challenge is engineering data quality at enormous scale: developing filtering systems, model-based quality scoring, dataset transformations, contamination detection, and tooling that allows researchers to understand and work with training corpora effectively.
We’d be especially interested in people who have worked on:
Experience with vLLM, SGLang, Docker, Kubernetes, infrastructure-as-code, experiment tracking or open-source NLP/data tooling would also be valuable.