Senior Software Engineer, Evaluation Flywheel – Autonomous Vehicles

Jobtailor

California (MO)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor seeks a senior ML evaluation engineer to own the eval flywheel strategy and architecture, shaping how road and simulation data become curated datasets and how metrics are measured against them. You will build tooling for rapid metric iteration and partner with senior engineers across test engineering, behavior planning, and infrastructure.

You will work directly with AI model developers to accelerate evaluation and ensure high-quality release processes as the system evolves.

Qualifications

  • 12+ years building software, with significant time in autonomous vehicles, robotics, or large-scale ML systems.
  • Deep experience evaluating ML or robotic systems: metric design, ground-truth and golden dataset curation, precision/recall methodology, and the data pipelines behind them.
  • Strong Python and data engineering skills for production-scale pipelines.
  • BS or MS in Computer Science, Robotics, or a related field (or equivalent experience).

Responsibilities

  • Own the eval flywheel strategy and architecture.
  • Set evaluation quality standards and maintain golden datasets.
  • Build tooling for rapid metric iteration and reporting.
  • Partner with senior engineers across test engineering, behavior planning, and infrastructure.

Skills

Machine Learning Evaluation
Golden Dataset Curation
Metric Performance Measurement
Python Programming
Data Engineering
Software Development
Robotics
Autonomous Vehicles
Large-Scale ML Systems
Ground-Truth Curation

Education

BS/MS in Computer Science or Robotics

Tools

Data Pipelines
Self-Serve Dataset Tools
Quality Reporting Tools

Job description

Owning the eval flywheel's strategy and architecture: how road and simulation driving data becomes curated golden datasets, how metrics are measured against them (precision/recall), and how those results earn lasting trust with the teams that depend on them.
Setting the standard for evaluation quality: golden dataset curation, versioning, and health; metric performance measurement; and release processes that keep results dependable as the system evolves.
Building the tooling that helps our metric developers iterate quickly: self-serve dataset pipelines, metric performance measurement, and quality reporting used every day by the team and our partners.
Partnering with senior engineers and leaders across test engineering, behavior planning, and infrastructure — setting expectations, working through trade-offs, and being the voice of evaluation quality in cross-team decisions.
Working directly with AI model developers so evaluation iteration speed becomes an advantage for the whole program, including our push into learned, VLM-based evaluation.

Requirements
  • 12+ years building software, with significant time in autonomous vehicles, robotics, or large-scale ML systems.
  • Deep experience evaluating ML or robotic systems: metric design, ground-truth and golden dataset curation, precision/recall methodology, and the data pipelines behind them.
  • Strong Python and data engineering skills for production-scale pipelines.
  • BS or MS in Computer Science, Robotics, or a related field (or equivalent experience).
Core Competencies

Demonstrates expertise in building and evaluating machine learning systems, with a focus on golden dataset curation, metric performance measurement, and data pipeline development. Strong collaboration with cross-functional teams to ensure evaluation quality and system reliability.

Highest-signal resume keywords
  • Machine Learning Evaluation
  • Golden Dataset Curation
  • Metric Performance Measurement
  • Python Programming
  • Data Engineering
ATS Optimization Keywords
Hard Skills
  • Machine Learning
  • Metric Design
  • Ground-Truth Curation
  • Precision/Recall Methodology
  • Data Pipeline Development
  • Software Development
  • Robotics
  • Autonomous Vehicles
  • Large-Scale ML Systems
  • Quality Reporting
Soft Skills
  • Collaboration
  • Communication
  • Strategic Thinking
  • Problem Solving
Industry Keywords
  • Autonomous Vehicles
  • Robotics
  • Machine Learning Systems
  • Evaluation Quality
  • Versioning
  • Release Processes
Tools & Technologies
  • Data Pipelines
  • Self-Serve Dataset Tools
  • Quality Reporting Tools
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles

NVIDIA • Washington

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles

NVIDIA • Georgia

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles

NVIDIA • Washington

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Software Engineer, Evaluation Flywheel - Autonomous Vehicles
Senior Software Engineer, Evaluation Flywheel - Autonomous Vehicles

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles

NVIDIA • North Carolina

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles
Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles

NVIDIA AI • Santa Clara (UT)

On-site
USD 224,000 - 357,000
Equity
Benefits
Principal Engineer, Tech Lead – Embodied AI, Off-Board Performance Evaluation
Principal Engineer, Tech Lead – Embodied AI, Off-Board Performance Evaluation

Jobtailor • Massachusetts

On-site
USD 180,000 - 260,000
Senior Manager, Data Science
Senior Manager, Data Science

Jobtailor • Marysville (OH)

On-site
USD 120,000 - 180,000
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000