Staff Engineer, Autonomous Driving Data Platform & Curation

CARIAD, Inc.

Mountain View (CA)

On-site

USD 162,000 - 234,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision
401k with employer match
Vacation and paid holidays
Vehicle lease program
Tuition reimbursement
Employee assistance program

Job summary

CARIAD, Inc. is seeking a Staff Engineer for Autonomous Driving Data Platform & Curation in Mountain View, CA.

This hands-on role owns end-to-end data from raw multimodal vehicle data to reliable, versioned, model-ready datasets, leading ingestion, schema, storage, and quality pipelines. You will design scalable data representations and pipelines for large multimodal data, collaborate across teams, and advance data-quality monitoring, labeling workflows, and reproducibility for imitation

Qualifications

  • 8+ years in applied ML or deep learning.
  • 4+ years in RL, CV, or AD/ADAS systems.
  • Strong dataset and storage foundation with schema design and versioning.
  • Hands-on with Parquet, Lance, Arrow; storage, partitioning, indexing.
  • Python and SQL with production engineering practices.
  • Experience building distributed data pipelines (Ray, Spark, Beam, etc.).
  • Familiarity with S3/GCS object storage and data governance.

Responsibilities

  • Design canonical representations for drives, scenarios, frames, and dataset manifests.
  • Build and maintain validated datasets >10M records on storage.
  • Evaluate Parquet, Lance, Arrow; own partitioning, schema evolution, lineage.
  • Optimize filtering, projection, joins, sampling, and data loading performance.
  • Measure ingestion throughput, query latency, and cost per usable sample.
  • Enable self-service discovery and data quality tagging for ML datasets.
  • Design automated data validation, dashboards, and observability.
  • Lead monitoring, root-cause analysis, and reliability improvements.

Skills

Production data pipelines
Large datasets
Parquet/Lance/Arrow
Python & SQL (prod practices)
Distributed data pipelines (Ray/Spark/
Cloud object storage (S3/GCS)
Autonomous driving data knowledge
Data quality & observability

Education

Master’s in Computer Science/Robotics/Engineering/Applied Math
PhD in related field (desirable)

Tools

Parquet
Lance
Apache Arrow
Spark
Ray
Beam
Airflow
Dagster

Job description

Role Summary

The Staff Engineer, Autonomous Driving Data Platform & Curation is a hands‑on staff‑level individual contributor who owns the path from raw multimodal vehicle data to reliable, versioned, and model-ready datasets. The ideal candidate combines production data or ML systems experience with practical knowledge of autonomous‑driving data and can lead ingestion, schema, storage, query, curation, quality, and delivery. This engineer designs and operates datasets containing more than 10 million records or samples, evaluates formats such as Parquet and Lance, and enables efficient filtering, slicing, random access, and sequential retrieval. The role also advances data‑quality monitoring, statistical and out‑of‑distribution detection, rule‑based and model‑based tagging, model‑in‑the‑loop and human‑in‑the‑loop labeling, and future data preparation for imitation learning and reinforcement learning.

Role Responsibilities
  • Design canonical representations for drives, scenarios, clips, frames, trajectories, sensor references, vehicle state, map context, labels, predictions, and dataset manifests.
  • Build and maintain validated, versioned datasets containing more than 10 million records or samples on local or cloud object storage.
  • Evaluate Parquet, Apache Arrow, Lance, and related technologies; own partitioning, indexing, file sizing, compaction, schema evolution, lineage, and reproducibility.
  • Optimize filtering, projection, joins, scenario slicing, random sampling, shuffling, sequential retrieval, and model data‑loading performance.
  • Measure and improve ingestion throughput, query latency, training throughput, storage utilization, reliability, and cost per usable sample.
  • Work effectively with synchronized camera and other sensor data, ego state, localization, calibration, coordinate frames, map context, control actions, clips, and trajectories.
  • Translate perception, planning, VLA, and evaluation needs into schemas, searchable attributes, scenario definitions, sampling strategies, and reproducible dataset splits.
  • Build reliable batch or distributed pipelines for ingestion, transformation, enrichment, validation, cataloging, and publication, including retries, backfills, idempotency, and observability.
  • Enable self‑service discovery and composition for lane keeping, lane changes, long‑tail scenarios, hard examples, balanced datasets, and leakage‑resistant train, validation, and test splits.
  • Define automated quality gates, dashboards, and alerts for completeness, validity, freshness, duplication, synchronization, calibration, corruption, label integrity, coverage, balance, and cost.
  • Apply statistics, sampling, and distribution comparisons to detect drift, anomalies, underrepresented conditions, and out‑of distribution data.
  • Use deterministic checks, heuristic rules, geometry and metadata queries, embeddings, VLMs, and learned models to assess quality and create searchable scenario tags.
  • Quarantine suspicious data and lead root‑cause analysis across collection, synchronization, schema, transformation, storage, annotation, sampling, and model‑consumer failures.
  • Design model‑in‑the‑loop and human‑in‑the‑loop labeling workflows with confidence thresholds, review routing, audit sampling, disagreement handling, and label provenance.
  • Close the loop from model failures and edge cases through selection, annotation, quality review, dataset publication, training, and evaluation.
  • Evaluate auto‑labeled and synthetic data using label quality, coverage, distributional impact, and downstream model performance.
  • Prepare replayable trajectory data for imitation learning and offline RL, including observations, actions, timestamps, policy versions, interventions, rewards, termination conditions, and alignment checks.
  • Serve as the staff‑level owner and escalation point for dataset architecture, storage and query performance, data quality, reliability, reproducibility, and cost.
  • Align vehicle collection, autonomy metadata, model interfaces, annotation, storage, governance, and training requirements across partner teams.
  • Communicate risks, trade‑offs, decisions, and recovery plans with clear evidence; maintain traceable schemas, data contracts, lineage, operating procedures, and incident records.
  • Mentor engineers and strengthen design reviews, testing, observability, reproducibility, cost awareness, and data engineering standards.
General Skills
  • Deep production experience with Parquet and/or Lance, including dataset migration, indexing, compaction, schema evolution, version maintenance, cloud storage, random access, or training‑loader optimization.
  • Experience with large‑scale multimodal autonomous‑driving data from cameras, LiDAR, radar, maps, localization, vehicle state, simulation, or other sensor‑rich robotic systems.
  • Experience building auto‑labeling, active‑learning, model‑in‑the‑loop, or human‑in‑the‑loop systems, including confidence calibration, annotation routing, review workflows, audit sampling, label agreement, and annotation QA.
  • Strong statistical knowledge for quality monitoring, distribution comparison, drift measurement, anomaly detection, and out‑of‑distribution detection.
  • Experience using rule‑based logic, geometry and metadata heuristics, embeddings, semantic search, VLMs, classifiers, or other learned methods to tag data, assess quality, mine edge cases, detect hard examples, or prioritize review.
  • Experience connecting dataset metrics, composition, label quality, drift, cost, and downstream model performance through dashboards and evaluation.
  • Experience optimizing distributed ML data loading, caching, sampling, preprocessing, and the data‑to‑GPU path.
  • Experience preparing sequential or trajectory datasets for imitation learning, offline reinforcement learning, preference or ranking data, reward modeling, or policy evaluation, including observations, actions, rewards, policy provenance, and replay semantics.
  • Familiarity with data governance, privacy, retention, redaction, access control, auditability, and security requirements for fleet‑collected sensor data.
Required Specialized Skills
  • Demonstrated ownership of production data pipelines and large datasets, including systems with more than 10 million records or samples or comparable multi‑terabyte scale.
  • Strong dataset and storage foundation, including schema design, columnar storage, partitioning, file and fragment sizing, compression, predicate pushdown, projection, indexing, compaction, schema evolution, versioning, lineage, and reproducibility.
  • Hands‑on experience with Parquet, Lance, Apache Arrow, or a comparable analytical dataset technology, with the ability to reason about format selection and benchmark storage, query, update, and model‑loading trade‑offs.
  • Strong Python and SQL skills with production software‑engineering practices, including testing, code review, version control, profiling, debugging, maintainable APIs, and operational documentation.
  • Experience building distributed or parallel data pipelines and orchestration using technologies such as Ray, Spark, Beam, Dask, Airflow, Dagster, or comparable systems.
  • Experience with local or cloud object storage, such as S3 or GCS, and practical understanding of throughput, request patterns, caching, data movement, reliability, access control, and cost.
  • Basic autonomous‑driving or robotics data knowledge, including temporal sensor data, camera and other sensor modalities, timestamps and synchronization, calibration, ego and vehicle state, coordinate frames, clips, trajectories, scenarios, and annotations.
  • Ability to design automated data validation and observability and translate ML training, sampling, evaluation, and failure‑analysis needs into reliable data products.
Desired Skills
  • Deep production experience with Parquet and/or Lance, including dataset migration, indexing, compaction, schema evolution, version maintenance, cloud storage, random access, or training‑loader optimization.
  • Experience with large‑scale multimodal autonomous‑driving data from cameras, LiDAR, radar, maps, localization, vehicle state, simulation, or other sensor‑rich robotic systems.
  • Experience building auto‑labeling, active‑learning, model‑in‑the‑loop, or human‑in‑the‑loop systems, including confidence calibration, annotation routing, review workflows, audit sampling, label agreement, and annotation QA.
  • Strong statistical knowledge for quality monitoring, distribution comparison, drift measurement, anomaly detection, and out‑of‑distribution detection.
  • Experience using rule‑based logic, geometry and metadata heuristics, embeddings, semantic search, VLMs, classifiers, or other learned methods to tag data, assess quality, mine edge cases, detect hard examples, or prioritize review.
  • Experience connecting dataset metrics, composition, label quality, drift, cost, and downstream model performance through dashboards and evaluation.
  • Experience optimizing distributed ML data loading, caching, sampling, preprocessing, and the data‑to‑GPU path.
  • Experience preparing sequential or trajectory datasets for imitation learning, offline reinforcement learning, preference or ranking data, reward modeling, or policy evaluation, including observations, actions, rewards, policy provenance, and replay semantics.
  • Familiarity with data governance, privacy, retention, redaction, access control, auditability, and security requirements for fleet‑collected sensor data.
Years Of Relevant Experience
  • 8+ years of experience in applied machine learning or deep learning
  • 4+ years of experience reinforcement learning, computer vision, or AD/ADAS systems.
  • Strong candidates with equivalent industry experience will be considered
Required Education

Master’s in Computer Science, Robotics, Electrical Engineering, Applied Mathematics, or a related field

Desired Education

PhD in Computer Science, Robotics, Electrical Engineering, Applied Mathematics, or a related field

Compensation

Salary range is dependent on factors such as geographical differentials, credentials or certifications, industry-based experience, qualification and training. In the city of Mountain View, CA, the salary range for this position is $161,710 - $234,325.

CARIAD, Inc. provides performance based merits and annual bonus along with a competitive benefits package. Benefits include medical, dental, vision, 401k with employer match and defined contribution plan, short and long term disability, basic life and AD&D insurance, employee assistance program, tuition reimbursement and student loan repayment plans, maternity and non-primary caregiver leave, adoption assistance, employee referral program and vacation and paid holidays. We also offer a unique vehicle lease program that covers registration and insurance fees.

CARIAD is an Equal Opportunity Employer. We welcome and encourage applicants from all backgrounds, and do not discriminate based on race, sex, age, disability, sexual orientation, national origin, religion, color, gender identity/expression, marital status, veteran status, or any other characteristics protected by applicable laws.

Employment with Cariad Inc. is contingent upon the successful completion of this screening process. We emphasize the importance of compliance with export control and sanctions laws as a fundamental aspect of our operations. Our company is dedicated to adhering to these regulations to ensure the lawful and ethical conduct of our business activities. Employment with our company is contingent on either verifying U.S. citizenship or U.S. lawful permanent resident status or obtaining any necessary license or confirming the availability of an applicable exemption or license exception. You, the applicant, will be required to answer certain questions for export control purposes, and that information will be reviewed by compliance personnel to ensure compliance with federal law. Cariad Inc. may choose not to apply for a license or use an applicable license exception (if available) for such individuals whose access to export‑controlled technology or software source code may require authorization and may decline to proceed with an applicant on that basis alone.

By submitting your application, you acknowledge and agree to participate in the export control and sanctions compliance screening process. Your cooperation in this matter is essential to our shared success and the integrity of our operations. Thank you for your understanding and commitment to upholding these important standards.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer, Autonomous Driving Data Platform & Curation New Mountain View, CA
Staff Engineer, Autonomous Driving Data Platform & Curation New Mountain View, CA

CARIAD, Inc. • Mountain View (CA)

Hybrid
USD 162,000 - 234,000
Medical insurance
Dental insurance
Vision insurance
+8
Staff Engineer, HIL Design, Hardware Integration & Test
Staff Engineer, HIL Design, Hardware Integration & Test

Cariad, Inc. • Mountain View (CA)

On-site
USD 162,000 - 234,000
Medical benefits
Dental benefits
Vision benefits
+9
Staff Engineer, HIL Design, Hardware Integration & Test Automation
Staff Engineer, HIL Design, Hardware Integration & Test Automation

Socket.dev • Mountain View (CA)

On-site
USD 161,710 - 234,325
Senior AI Data Pipeline Engineer (Autonomous Driving)
Senior AI Data Pipeline Engineer (Autonomous Driving)

42dot Inc. • Sunnyvale (CA)

On-site
USD 133,000 - 254,000
Robotics Engineer II
Robotics Engineer II

Robotgifs.com • Ann Arbor (MI)

On-site
USD 130,000 - 155,000
Comprehensive healthcare
Generous paid parental leave
Flexible vacation
+2
Robotics Engineer II
Robotics Engineer II

May Mobility • United States

On-site
USD 130,000 - 155,000
Healthcare benefits
HSA / FSA
Retirement plan
+3
Robotics Engineer II
Robotics Engineer II

Maymobility • Ann Arbor (MI)

On-site
USD 130,000 - 155,000
Comprehensive healthcare suite
Rich retirement benefits
Flexible vacation policy
Robotics Engineer II
Robotics Engineer II

Voiceflow • Ann Arbor (MI)

On-site
USD 130,000 - 155,000
Comprehensive healthcare
Rich retirement benefits
Generous parental leave
+1
Software Engineer II - Data Platform
Software Engineer II - Data Platform

May Mobility • Ann Arbor (MI)

On-site
USD 110,000 - 157,000
Comprehensive healthcare
Employer retirement match
Generous parental leave
+3
Robotics Engineer II New Ann Arbor, MI - Onsite
Robotics Engineer II New Ann Arbor, MI - Onsite

May Mobility, Inc. • Ann Arbor (MI)

On-site
USD 130,000 - 155,000
Healthcare benefits
401(k) with employer match
Paid parental leave
+1