ML Data Engineer: Multimodal Data Pipelines + Stock Options

Sesame

Bellevue (WA)

On-site

USD 170,000 - 280,000

Full time

24 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

401(k) match
Health, vision, dental coverage
Unlimited PTO
Flexible spending account
EAP
Stock options

Job summary

Sesame is seeking a Data Engineer to build and maintain data pipelines that feed our AI models, working closely with ML engineers and researchers to ensure data is clean, versioned, and ready for training and evaluation.

You will design systems handling multimodal data, including conversations and sensor signals, turning raw data into trusted datasets for model development, with a strong focus on data quality and governance.

Qualifications

  • 5+ years in data engineering, with ML/AI support experience.
  • Strong SQL and Python skills used daily.
  • Experience building and operating ETL/ELT pipelines at scale.
  • Experience with workflow orchestration systems (Airflow, Dagster, Prefect).
  • Hands-on experience with ML data workflows: training data pipelines, dataset versioning, or eval data.
  • Understanding of how ML teams work and data quality impact on models.
  • Comfort with unstructured data (audio, text, JSON logs).
  • Strong communication to bridge data systems and model needs.

Responsibilities

  • Design and build production data pipelines for conversational, voice, and multimodal data used in model training and evaluation.
  • Collaborate with ML engineers to define data requirements for new models and experiments.
  • Maintain dataset versioning, lineage tracking, and reproducibility for training runs.
  • Develop data quality frameworks: schema validation, drift detection, coverage monitoring.
  • Optimize large-scale data processing for cost and performance across cloud infra.
  • Create tooling for ML engineers to discover, explore, and request data.
  • Define data governance and privacy standards for sensitive data.
  • Contribute to architecture decisions for Sesame's data platform as volume grows.

Skills

SQL
Python
ETL pipelines
Workflow orchestration
ML data workflows
Data governance
Unstructured data
Communication

Tools

Airflow
Dagster
Prefect
Ray
Spark
Kubernetes

Job description

Sesame is seeking a Data Engineer to build and maintain data pipelines that feed our AI models, working closely with ML engineers and researchers to ensure data is clean, versioned, and ready for training and evaluation.

You will design systems handling multimodal data, including conversations and sensor signals, turning raw data into trusted datasets for model development, with a strong focus on data quality and governance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Data Engineer - Multimodal AI Pipelines & Datasets
ML Data Engineer - Multimodal AI Pipelines & Datasets

Sesame • New York (NY)

On-site
USD 170,000 - 280,000
401(k) match
Employer-paid health, vision, dental
Unlimited PTO
+3
ML Data Engineer: Scalable Pipelines & Data Quality
ML Data Engineer: Scalable Pipelines & Data Quality

Sesame • San Francisco (CA)

On-site
USD 120,000 - 160,000
401(k) max employer match: 3.5% of compensation
100% employer-paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
Senior ML Data Engineer — Scalable Multimodal Pipelines
Senior ML Data Engineer — Scalable Multimodal Pipelines

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Multimodal Data Engineer: End-to-End Pipelines + Equity
Multimodal Data Engineer: End-to-End Pipelines + Equity

ABAKA AI • Mountain View (CA)

On-site
USD 120,000 - 225,000
Equity in Abaka AI
Comprehensive benefits: health, dental
Vision coverage
+1
ML Data Engineer — Multimodal AI Pipelines
ML Data Engineer — Multimodal AI Pipelines

Veeda AI • Seattle (WA)

On-site
USD 140,000 - 190,000
Senior Data Engineer - ML Ops & Multimodal Data
Senior Data Engineer - ML Ops & Multimodal Data

tavus • United States

Remote
USD 120,000 - 190,000
Multimodal ML Engineer — Production Systems
Multimodal ML Engineer — Production Systems

Jack • San Francisco (CA)

On-site
USD 150,000 - 350,000
ML Data & Platform Engineer: Pipelines, MLOps & Production
ML Data & Platform Engineer: Pipelines, MLOps & Production

Speechmatics • Cambridge (MA)

Hybrid
USD 150,000 - 190,000
Private Medical
Dental for you and family
Pension/401K matching
+4
Data Engineer, Machine Learning
Data Engineer, Machine Learning

Sesame • San Francisco (CA)

On-site
USD 120,000 - 160,000
401(k) max employer match: 3.5% of compensation
100% employer-paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
ML Engineer: Multimodal Data & Production Systems
ML Engineer: Multimodal Data & Production Systems

Sieve, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000