Senior Test Engineer - Metrics Quality

Avride

Austin (TX)

On-site

USD 110,000 - 160,000

Full time

48 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Avride is seeking a QA engineer to own the quality of the metrics that drive decisions in our autonomous driving software. You’ll assess whether metrics can be trusted before release decisions, reading implementations and understanding product goals.

You’ll build test data, map case spaces, and uncover silent failures where metrics drift or report nothing, ensuring metrics truly reflect system behavior across scenarios.

Qualifications

  • 6+ years in software testing or test engineering in data-heavy or analytical systems.
  • Strong Python and SQL skills; you write analysis scripts and independent reference computations.
  • Statistical literacy: distributions, variance, sample size, aggregation traps, and confidence to call noise on differences.
  • Product thinking: ask what a number is for before you test it; testing against specs is just the start.
  • You’ll regularly push back when a metric behaves unexpectedly, with evidence.
  • Fluent written English; findings read by cross-functional teams.
  • Proficiency with LLMs as a working tool for analysis, data generation, and cross-checking reasoning.

Responsibilities

  • Own the acceptance of new and changed metrics before they’re used to ship; decide what to count and what not to.
  • Build the case space for each metric, covering all situations including silent ones.
  • Hunt silent failures where metrics return nothing when they should report something.
  • Read the implementation against the case space to find uncovered real-world scenarios.
  • Build test data as needed, using routes that best fit the case at hand.
  • Present findings to engineers with concrete examples, scale, and impact.
  • Think from the reader’s perspective; describe degradation across the whole set, not just per scene.
  • Own the release cycle for metrics, qualifying releases and maintaining regression coverage.

Skills

Python
SQL
Statistical analysis
LLMs usage
Written communication
Test automation
Data-heavy systems

Tools

ClickHouse

Job description

About The Team

Every decision Avride makes about its autonomous driving software merge or revert, ship or hold, this approach or that one — rests on a metric computed over a set of recorded scenes. Our QA organization is what stands between a metric and a decision made on a metric that was quietly wrong.

About The Team

Every decision Avride makes about its autonomous driving software merge or revert, ship or hold, this approach or that one — rests on a metric computed over a set of recorded scenes. Our QA organization is what stands between a metric and a decision made on a metric that was quietly wrong.

About The Role

You will own the quality of the metrics themselves. Our data scientists build them; you decide whether they can be trusted, before anyone starts making release decisions with them. This is a specific and underrated craft. A metric can be computed correctly, pass every unit test, and still be the wrong number for the decision it is meant to support. Finding that takes reading the implementation, understanding the product, and knowing what the people who rely on the number are actually trying to learn from it. The hard half of the job is not checking that a metric fires when it should. It is finding the places where it stays silent and should not have — the failures that produce no number at all, and therefore no complaint, until someone ships on the strength of them.

What you’ll do
  • Own the acceptance of new and changed metrics. Before a metric is used to decide whether a change is safe to ship, you decide whether it can be: what it should count, what it should not, and whether the implementation agrees with either.
  • Build the case space. For each metric, work out the full set of situations it has to handle — including all the ones where it must stay silent — and keep that set current as the technology and the operating environment change.
  • Hunt the silent failures. A metric that returns a surprising number gets noticed. A metric that returns nothing where it should have returned something does not, and that is the class of defect you are here to find.
  • Read the implementation against the case space. Not for code quality — for the real situations it does not handle, and will therefore never report.
  • Build the test data the job needs. The situations a metric has to handle are rarely all sitting in the data already. Get them by whatever route is cheapest for the case at hand, and keep looking for better routes than the ones we use today.
  • Make the case for a fix. Take findings to the engineer who owns the metric with concrete examples and a sense of scale: how often it is wrong, and what decisions that changes.
  • Think like the people who read the numbers. A metric that is right on every individual scene can still fail to reveal a degradation across the whole set. Say so before anyone builds a release gate on it.
  • Own the release cycle for metrics. Metrics ship on their own cadence. You qualify each release, run and maintain the regression that catches silent changes in what a metric means rather than only in what it returns, and make that cycle faster and cheaper as the number of metrics grows.
What you’ll need
  • 6+ years in software testing or test engineering, with real depth in data-heavy or analytical systems.
  • Strong Python and SQL. You write analysis scripts and independent reference computations as a matter of routine, not as an exception.
  • Statistical literacy. Distributions, variance, sample size, aggregation traps, and the confidence to say "this difference is noise".
  • Product thinking. You ask what a number is for before you test it. Testing against the letter of a specification is the starting point of this job, not the substance of it.
  • Conviction that survives a "that's how it works." You will regularly be the one telling a colleague that their metric does not behave the way they expect. That is a normal part of the work here, and it goes well when you bring evidence and see the conversation through rather than letting the finding quietly drop.
  • Fluent use of LLMs as a working tool — analysis, test-data generation, reading unfamiliar code, and cross-checking your own reasoning.
  • Clear written English. Your findings are read by people who will act on them.
Nice to have
  • Autonomous vehicles, robotics, or evaluation of machine learning systems.
  • ClickHouse or a similar analytical database.
  • Experience testing evaluation pipelines or A/B testing infrastructure.
  • Experience designing annotation tasks for human reviewers, and checking that their answers agree.
  • Knowledge of US road rules and driver behaviour.

Candidates are required to be authorized to work in the U.S. The employer is not offering relocation, sponsorship, and remote work options are not available.

Avride is an equal opportunity employer and committed to providing reasonable accommodations to qualified applicants and employees with disabilities to ensure they have equal access to employment opportunities. Avride complies with the Americans with Disabilities Act (ADA), if you need a reasonable accommodation to assist with the application or hiring process, or to perform the essential functions of a job, please email jobs@avride.ai.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Test Engineer - Scenario Coverage & Evaluation Data Sets
Lead Test Engineer - Scenario Coverage & Evaluation Data Sets

Avride • Austin (TX)

On-site
USD 140,000 - 180,000
Lead Test Engineer - Scenario Coverage & Evaluation Datasets
Lead Test Engineer - Scenario Coverage & Evaluation Datasets

Avride • Austin (TX)

On-site
USD 140,000 - 190,000
Data Scientist – AV Metrics & Evaluation Analytics
Data Scientist – AV Metrics & Evaluation Analytics

Avride • Austin (TX)

On-site
USD 90,000 - 130,000
Lead Test Engineer AV Safety Assurance
Lead Test Engineer AV Safety Assurance

Avride • Austin (TX)

On-site
USD 150,000 - 210,000
Data Scientist – AV Metrics & Evaluation Analytics
Data Scientist – AV Metrics & Evaluation Analytics

Avride Inc. • Austin (TX)

On-site
USD 120,000 - 160,000
Equal opportunity employer
Backend Engineer - Metrics Infrastructure
Backend Engineer - Metrics Infrastructure

Avride • Austin (TX)

On-site
USD 90,000 - 135,000
Medical insurance
Vision insurance
401(k)
+1
Data Quality Analyst- Autonomous Vehicles
Data Quality Analyst- Autonomous Vehicles

Avride • Austin (TX)

On-site
USD 95,000 - 130,000
Senior Data Scientist - Metrics
Senior Data Scientist - Metrics

Pantera Capital • Ann Arbor (MI)

On-site
USD 163,000 - 241,000
Comprehensive healthcare suite
Health Savings Accounts
Rich retirement benefits
+3
Senior Data Scientist - Metrics
Senior Data Scientist - Metrics

May Mobility • Ann Arbor (MI)

On-site
USD 163,000 - 241,000
Comprehensive healthcare suite
Flexible vacation policy
Rich retirement benefits
+2
Senior Data Scientist - Metrics Ann Arbor, MI
Senior Data Scientist - Metrics Ann Arbor, MI

May Mobility • Ann Arbor (MI)

On-site
USD 163,000 - 241,000
Comprehensive healthcare suite
Rich retirement benefits
Flexible vacation policy
+1