ML Data Engineer - Multimodal AI Pipelines & Datasets

Sesame

New York (NY)

On-site

USD 170,000 - 280,000

Full time

10 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

401(k) match
Employer-paid health, vision, dental
Unlimited PTO
FSA with employer match
EAP
Stock options

Job summary

Sesame is hiring a Data Engineer to design and maintain data pipelines for its AI models. You will work closely with ML engineers to ensure clean, versioned, and well-documented datasets for training, evaluation, and deployment.

The role is deeply technical and infrastructure-focused, embedding with ML teams to accelerate the full model development lifecycle from data collection and labeling through training and evaluation.

Qualifications

  • 5+ years in data engineering, with meaningful experience supporting ML or AI teams specifically.
  • Strong SQL and Python skills — you'll use both daily.
  • Experience building and operating ETL/ELT pipelines at scale using modern data platforms and tooling.
  • Experience with workflow orchestration systems such as Airflow, Dagster, or Prefect.
  • Hands-on experience with ML data workflows: training data pipelines, dataset versioning, data labeling pipelines, or model evaluation data.
  • A solid understanding of how ML teams work and what makes a good training dataset.

Responsibilities

  • Design and build production data pipelines that prepare conversational, voice, and multimodal data for model training and evaluation.
  • Partner directly with ML engineers to understand data requirements for new models and experiments, and deliver datasets that meet those needs.
  • Build and maintain infrastructure for dataset versioning, lineage tracking, and reproducibility — so any training run can be traced back to its exact data.
  • Develop data quality frameworks that catch issues before they become model quality issues: schema validation, drift detection, and coverage monitoring.
  • Optimise large-scale data processing for cost and performance across Sesame's cloud infrastructure.
  • Build tooling that makes it easy for ML engineers and researchers to discover, explore, and request data independently.
  • Define and enforce data governance and privacy standards, particularly around sensitive conversational and voice data.
  • Contribute to architecture decisions around Sesame's broader data platform as the team and data volume grow.

Skills

Data engineering
SQL
Python
ML data workflows
Unstructured data

Tools

Airflow
Dagster
Prefect
Spark
Kubernetes

Job description

Sesame is hiring a Data Engineer to design and maintain data pipelines for its AI models. You will work closely with ML engineers to ensure clean, versioned, and well-documented datasets for training, evaluation, and deployment.

The role is deeply technical and infrastructure-focused, embedding with ML teams to accelerate the full model development lifecycle from data collection and labeling through training and evaluation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Data Engineer: Multimodal Data Pipelines + Stock Options
ML Data Engineer: Multimodal Data Pipelines + Stock Options

Sesame • Bellevue (WA)

On-site
USD 170,000 - 280,000
401(k) match
Health, vision, dental coverage
Unlimited PTO
+3
ML Data Engineer: Scalable Pipelines & Data Quality
ML Data Engineer: Scalable Pipelines & Data Quality

Sesame • San Francisco (CA)

On-site
USD 120,000 - 160,000
401(k) max employer match: 3.5% of compensation
100% employer-paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
Remote Data Engineer for AI/ML Pipelines
Remote Data Engineer for AI/ML Pipelines

Equiliem • United States

Remote
USD 171,924,000 - 186,252,000
Medical Insurance
Vision & Dental Insurance
Life Insurance
+4
ML Data Engineer — Multimodal AI Pipelines
ML Data Engineer — Multimodal AI Pipelines

Veeda AI • Seattle (WA)

On-site
USD 140,000 - 190,000
ML Data Infrastructure Engineer: Video & Sensor Pipelines
ML Data Infrastructure Engineer: Video & Sensor Pipelines

HavocAI • Town of Providence (NY)

On-site
USD 140,000 - 180,000
Health insurance
Life Insurance
401k Matching
+6
Lead Data Engineer for Multimodal ML Pipelines
Lead Data Engineer for Multimodal ML Pipelines

Sarah Smith Fund • United States

Hybrid
USD 150,000 - 230,000
Competitive salary
Equity
Health benefits
+3
Senior ML Engineer: LLM Data Pipelines & Datasets
Senior ML Engineer: LLM Data Pipelines & Datasets

Cisco • San Jose (CA)

Hybrid
USD 203,000 - 259,000
Data Engineer, Machine Learning
Data Engineer, Machine Learning

Sesame • San Francisco (CA)

On-site
USD 120,000 - 160,000
401(k) max employer match: 3.5% of compensation
100% employer-paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
ML Data Infra Engineer - Build Autonomous Data Pipelines
ML Data Infra Engineer - Build Autonomous Data Pipelines

HavocAI • United States

On-site
USD 120,000 - 190,000
Health, Dental and Vision Insurance (U
Life Insurance (Employer Paid)
401k Matching
+5
Senior ML Data Engineer — Scalable Multimodal Pipelines
Senior ML Data Engineer — Scalable Multimodal Pipelines

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000