Member of Technical Staff, Data & ML Infrastructure for Video Models

Cantina Labs

Greater London

On-site

GBP 150,138 - 195,179

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical insurance
Paid time off
Parental leave
401(k) retirement plan
Lifestyle spending account
Office lunch and snacks
Equity

Job summary

Cantina Labs is hiring a Member of Technical Staff to build and scale data pipelines for our large video generation models. You will own annotation workflows, dataset curation, and preprocessing tools that directly improve model quality.

You’ll collaborate with research and engineering to turn experiments into scalable systems, ensuring data is clean, ready for training, and delivered efficiently within our Kubernetes-based infrastructure.

Qualifications

  • 3+ years of experience in machine learning, data pipelines or related engineering roles on large-scale multimodal or video systems.
  • Strong programming skills in Python and building reliable data processing pipelines for ML workflows.
  • Hands-on experience preparing training data including parsing, filtering, and dataset curation using cloud platforms.

Responsibilities

  • Build and maintain data pipelines for large video generation models including ingestion, parsing, filtering, preprocessing, and dataset curation at scale.
  • Design and run annotation workflows across MTurk, Prolific, and Mechanical Turk with quality control.
  • Train and evaluate smaller models used for data filtering and preprocessing in the ML pipeline.
  • Partner with research and engineering teams to scale experimental workflows into repeatable systems.
  • Own data quality across the pipeline and continuously improve tooling and processes.
  • Build internal tools to prepare datasets and monitor outputs for model development.
  • Drive large pipeline projects from inception to completion, including dataset upgrades.
  • Work within a Kubernetes-based training infra to ensure data is prepared and delivered to training clusters.
  • Profile and optimize inference scripts used in preprocessing for time and cost efficiency.

Skills

Python
Data pipelines
ML preprocessing
Kubernetes
MTurk/Prolific workflows
Cloud computing
PyTorch
Cross-functional collaboration
Large-scale data handling

Education

CS/ML degree

Tools

AWS S3
DynamoDB
RunPod

Job description

About Cantina

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

About Cantina

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About The Role

We are looking for a new Member of Technical Staff to build and scale the data pipelines behind our large video generation models. This role is focused on collecting large amounts of relevant video data, preparing high-quality training samples, and developing robust preprocessing, filtering, and parsing workflows. You'll orchestrate annotation pipelines across platforms such as MTurk and own the full lifecycle of training data, from raw ingestion to clean, model-ready samples that directly drive quality improvements. This role sits at the intersection of data engineering and ML research, making it central to how we turn messy real-world data into the fuel that moves our models forward.

What You’ll Do
  • Build and maintain data pipelines for large video generation models, including data ingestion, parsing, filtering, preprocessing, and dataset curation at scale, using tools such as AWS S3 and DynamoDB.
  • Design and run annotation workflows across platforms such as MTurk, Prolific, and Mechanical Turk, including task design, quality control, and label validation.
  • Train, evaluate, and improve smaller supporting models used for data filtering, quality assessment, preprocessing, or other parts of the ML pipeline.
  • Partner closely with research and engineering teams to turn experimental workflows into scalable, repeatable systems that support model training and evaluation.
  • Own data quality across the pipeline by identifying bottlenecks, failure modes, and low-quality sources, and continuously improving tooling and processes.
  • Build internal tools and automation that make it easier to prepare datasets, launch annotation jobs, monitor outputs, and support model development end to end.
  • Drive larger pipeline projects from start to finish, such as new dataset creation efforts or upgrades to labeling and preprocessing infrastructure.
  • Work within a Kubernetes-based training infrastructure, ensuring datasets are properly prepared, formatted, and delivered to training clusters.
  • Profile and optimize research model inference scripts used in preprocessing steps, ensuring that model-driven filtering and transformation stages run within practical time and cost constraints when applied to large-scale raw data.
What You’ll Bring
  • 3+ years of experience in machine learning, applied ML, data pipelines, or related engineering roles, ideally working on large-scale multimodal, video, or vision-based systems.
  • Strong programming skills in Python and solid experience building reliable data processing and preprocessing pipelines for ML workflows.
  • Hands-on experience preparing training data for ML models, including parsing, filtering, dataset curation, quality control, and large-scale data handling using tools such as AWS S3 and DynamoDB.
  • Familiarity with annotation and labeling workflows, including task design, vendor or crowd-platform orchestration such as MTurk or Prolific, and methods for ensuring label quality.
  • Experience working with Kubernetes for orchestrating distributed workloads, including data preprocessing, pipeline execution, and dataset delivery to training clusters.
  • Comfort working across cloud and on-demand compute environments such as AWS and RunPod, with the ability to port and optimize pipelines across infrastructure.
  • Familiarity with distributed data processing frameworks and experience designing systems that operate reliably at scale across many nodes or workers.
  • Working knowledge of PyTorch and the broader deep learning stack, with the ability to read, debug, and optimize research model inference code for use in production preprocessing pipelines.
  • Ability to work cross-functionally with research and engineering teams and translate experimental ideas into robust, scalable systems.
  • Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Engineering, Mathematics, or a related technical field; experience in generative video, computer vision, or multimodal ML is strongly preferred.
  • Bonus: Experience training, evaluating, or fine-tuning smaller ML models used for classification, filtering, ranking, quality assessment, or other supporting tasks in an ML pipeline.
Compensation

The anticipated annual base salary range for this role is between $200,000-$260,000 (€170,000-€225,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits For U.S.-based Roles
  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
    • 15 PTO days
    • 10 sick days
    • 15 company holidays
    • 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, TTS
Machine Learning Engineer, TTS

Cantina Labs • Greater London

On-site
GBP 148,000 - 164,000
Equity
Health insurance
PTO 42 days
+3
Staff Data & ML Infra Engineer — Scale Video Pipelines
Staff Data & ML Infra Engineer — Scale Video Pipelines

Cantina Labs • Greater London

On-site
GBP 150,000 - 196,000
Medical insurance
Paid time off
Parental leave
+4
Senior Member of Technical Staff, Synthetic Data
Senior Member of Technical Staff, Synthetic Data

Cohere • Greater London

Hybrid
GBP 70,000 - 90,000
Inclusive culture and work environment
Work with a cutting-edge AI team
Weekly lunch stipend
+5
Senior Member of Technical Staff, Safety and Security for Agents
Senior Member of Technical Staff, Safety and Security for Agents

Cohere • City of Edinburgh

Hybrid
GBP 70,000 - 90,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+4
ML Engineer
ML Engineer

VEED • Greater London

Hybrid
GBP 76,000 - 120,000
Unlimited paid holidays
IT Equipment programme
Mental health benefit
Senior Member of Technical Staff, Safety and Security for Agents
Senior Member of Technical Staff, Safety and Security for Agents

Cohere • Greater London

Hybrid
GBP 90,000 - 150,000
Open and inclusive culture
Work with cutting-edge AI research team
Weekly lunch stipend
+5
Senior Research Data Engineer
Senior Research Data Engineer

Canva • Greater London

On-site
GBP 90,000 - 140,000
Member of Technical Staff, Pre-Training Data
Member of Technical Staff, Pre-Training Data

Cohere • City Of London

On-site
GBP 60,000 - 90,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+1
Senior Data Scientist
Senior Data Scientist

Carwow Group • Greater London

Hybrid
GBP 90,000 - 130,000
Hybrid working
Competitive salary
Share options
+3
Research Engineer in Data
Research Engineer in Data

Synthesia • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive compensation
Hybrid work setting
25 days of annual leave
+1