ML Data Engineer - End-to-End & Large-Scale Pipelines

Veeda Innovation

California (MO)

Hybrid

USD 130,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Veeda AI seeks a Member of Technical Staff - ML Data to design and validate data-centric ML methods. You will own the full lifecycle from problem formulation to deployment, building scalable pipelines for real-world datasets and synthetic data generation.

Strong Python and PyTorch expertise is required, along with a track record of rigorous experimental validation and reproducible software engineering practices. Join a fast-moving team tackling physical AI challenges.

Qualifications

  • Master’s or Ph.D. in Computer Science, Engineering, or a related technical field, or equivalent hands-on experience.
  • Demonstrated ability to develop original ML methods, evidenced by peer-reviewed publications or substantial research contributions with rigorous experimental validation.
  • Experience owning the full lifecycle of an ML method: designing and implementing the approach, applying it to large-scale data, evaluating results, and improving it through successive iterations.
  • Strong Python and PyTorch skills, with experience training, adapting, and evaluating machine learning models.
  • Ability to design controlled experiments, establish meaningful metrics, and analyze errors to guide improvements.
  • Strong software engineering skills, with an emphasis on reproducibility, reliability, and maintainable code.

Responsibilities

  • ML Methods for Data: Design, develop, and validate ML methods for data selection, enrichment, annotation, and quality assessment.
  • End-to-End Ownership: Own the full lifecycle, from problem formulation through deployment and continuous improvement.
  • Large-Scale Application: Build reliable workflows for processing large-scale real-world datasets and scalable pipelines for generating synthetic data.
  • Evaluation & Experimentation: Measure data quality and assess its impact on model performance.
  • Iterative Improvement: Use failure analysis and feedback to improve data-processing methods.
  • Annotation: Produce and evaluate labels such as captions, camera poses, depth maps, and segmentation masks.
  • Research Collaboration: Partner with researchers and engineers to develop effective data solutions.

Skills

Python
PyTorch
Experimentation
Software engineering

Education

Master's degree in CS/engineering
Ph.D. in related field

Tools

Docker
Linux

Job description

Veeda AI seeks a Member of Technical Staff - ML Data to design and validate data-centric ML methods. You will own the full lifecycle from problem formulation to deployment, building scalable pipelines for real-world datasets and synthetic data generation.

Strong Python and PyTorch expertise is required, along with a track record of rigorous experimental validation and reproducible software engineering practices. Join a fast-moving team tackling physical AI challenges.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Data Engineer - End-to-End & Large-Scale Pipelines
Staff ML Data Engineer - End-to-End & Large-Scale Pipelines

Veeda • California (MO)

On-site
USD 140,000 - 180,000
Member of Technical Staff - Data
Member of Technical Staff - Data

Veeda Innovation • California (MO)

Hybrid
USD 130,000 - 180,000
Member of Technical Staff - ML Data
Member of Technical Staff - ML Data

Veeda • California (MO)

On-site
USD 140,000 - 180,000
ML Systems Engineer — Performance & Scale
ML Systems Engineer — Performance & Scale

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000
Remote ML Data Engineer for Large-Scale AI Pipelines
Remote ML Data Engineer for Large-Scale AI Pipelines

Bright-Vision-Technologies • United States

Remote
USD 100,000 - 150,000
ML Data Engineer - DataOps & Pipelines
ML Data Engineer - DataOps & Pipelines

X Development, LLC • Mountain View (CA)

On-site
USD 166,000 - 244,000
Bonus
Equity
Benefits
Remote ML Data Engineer - Scale-Power Data Pipelines
Remote ML Data Engineer - Scale-Power Data Pipelines

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
AI Data Engineer — Build Scalable Cloud Data Pipelines for ML
AI Data Engineer — Build Scalable Cloud Data Pipelines for ML

AVP VIGILANT TECHNOLOGY PVT LTD • New York (NY)

On-site
USD 115,000 - 195,000
Competitive salary
Health benefits
401(k)
+4
Remote AI Data Engineer for Petabyte-Scale ML Pipelines
Remote AI Data Engineer for Petabyte-Scale ML Pipelines

Socket.dev • Ann Arbor (MI)

On-site
USD 80,000 - 100,000
ML Data Infra Engineer - Build Autonomous Data Pipelines
ML Data Infra Engineer - Build Autonomous Data Pipelines

HavocAI • United States

On-site
USD 120,000 - 190,000
Health, Dental and Vision Insurance (U
Life Insurance (Employer Paid)
401k Matching
+5