Get more replies from employers
Send a job-specific resume in minutes.
Jobtailor seeks a senior ML evaluation engineer to own the eval flywheel strategy and architecture, shaping how road and simulation data become curated datasets and how metrics are measured against them. You will build tooling for rapid metric iteration and partner with senior engineers across test engineering, behavior planning, and infrastructure.
You will work directly with AI model developers to accelerate evaluation and ensure high-quality release processes as the system evolves.
Owning the eval flywheel's strategy and architecture: how road and simulation driving data becomes curated golden datasets, how metrics are measured against them (precision/recall), and how those results earn lasting trust with the teams that depend on them.
Setting the standard for evaluation quality: golden dataset curation, versioning, and health; metric performance measurement; and release processes that keep results dependable as the system evolves.
Building the tooling that helps our metric developers iterate quickly: self-serve dataset pipelines, metric performance measurement, and quality reporting used every day by the team and our partners.
Partnering with senior engineers and leaders across test engineering, behavior planning, and infrastructure — setting expectations, working through trade-offs, and being the voice of evaluation quality in cross-team decisions.
Working directly with AI model developers so evaluation iteration speed becomes an advantage for the whole program, including our push into learned, VLM-based evaluation.
Demonstrates expertise in building and evaluating machine learning systems, with a focus on golden dataset curation, metric performance measurement, and data pipeline development. Strong collaboration with cross-functional teams to ensure evaluation quality and system reliability.