Technical Lead - Autonomy Evaluation

Atoms

San Francisco (CA)

On-site

USD 185,000 - 242,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
Disability Insurance
Life Insurance
401(k) Plan
Equity
Paid Holidays
Parental Leave
Commuter Benefit
Team lunch

Job summary

Atoms is seeking a Technical Lead for Autonomy Evaluation to own metric definitions and the log replay/evaluation platform for autonomous systems. You will guide evaluation from logs to scored test cases and ensure release decisions are well-supported by data.

You will lead onset of regression benchmarking, data mining for rare events, and scalable execution across GPU clusters, influencing the team and processes.

Qualifications

  • Bachelor's degree in Computer Science, Robotics, Electrical Engineering, Statistics, or related field.
  • At least 8+ years of relevant experience.
  • Experience as a technical lead or engineering manager for evaluation/testing.
  • Ownership of offline/system-level evaluation for autonomous systems or large-scale ML models.
  • Experience designing safety/behavior metrics for autonomous systems.
  • Hands-on experience building log replay at scale and scoring against ground truth.

Responsibilities

  • Own safety and behavior metrics used to judge releases.
  • Own the log replay and scoring pipeline for evaluation.
  • Benchmark software against previous releases and manage regression detection.
  • Design and execute evaluation data mining across log corpus.
  • Ensure deterministic, reproducible evaluation and track cost per replay-hour.
  • Set engineering practices for metrics, test design, code review, and reporting.

Skills

Leadership
Metrics design
Log replay
Benchmarking
Statistics
Autonomous systems
Python
C++ reading
Debugging
Communication

Education

Bachelor's degree in CS/Robotics/EE/Statistics
8+ years experience
Lead/evaluation experience

Tools

Python
C++

Job description

Who We Are

Atoms is building the machines that power the next era of progress.

Who We Are

Atoms is building the machines that power the next era of progress.

Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are moved, and industries are run, remains far less intelligent, far less efficient, and far more constrained. We’re changing that.

Atoms builds Physical AI - real-world robots for the industries that move civilization forward, starting with food, mining, and transport. Our systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive.

This work requires more than robotics. It requires deep integration across hardware, software, AI, operations, manufacturing, and real estate. We don’t just build machines in a lab. We deploy them into real environments, operate them, learn from them, and improve them until they work at scale.

We are roboticists, engineers, operators, and builders. We believe the next great technology companies will not only transform information, but the physical systems that shape everyday life.

If you want to work on hard problems with real-world impact, join us.

About The Role

We are seeking a Technical Lead for Autonomy Evaluation to own the metrics our autonomy releases are judged on and the log replay and evaluation platform that computes them. In this role, you will take evaluation from recorded logs, to scored regression suites running on every candidate release, to the report a release passes before it reaches a vehicle, developing how replay, scoring, metrics, and test coverage come together into one platform engineers use every day. You will do this across on-road vehicles, and you will build the evaluation infrastructure that lets each new software release, sensor configuration, and platform be assessed seamlessly.

What You’ll Do
  • Metrics. Own the safety and behavior metrics a release is judged on, including collision and near‑miss measures, trajectory agreement against ground truth, and comfort, and the go/no‑go release criteria built on them.
  • Log replay and scoring. Own the pipeline that converts recorded logs into scored test cases, covering open-loop and closed-loop evaluation of perception, localization, and planning against ground truth.
  • Regression and benchmarking. Own the system that benchmarks new software against old across the log corpus for every candidate release, including evaluation dataset design, regression detection, and triage of a regressed case to an owner.
  • Test coverage and data mining. Own how the corpus is mined for the long tail of rare and important events, including sampling strategy, selection of logs by expected value per replay‑hour, and coverage of scenario classes the corpus does not yet contain.
  • Execution at scale. Own deterministic and reproducible execution of evaluation, run identity, traceability of every result to a software version, and cost per replay‑hour as a tracked number.
  • Engineering practices. Set the practices for metric definitions, test design, code review, and the write‑up of evaluation results.
  • Technical standard. Mentor the engineers who join the function, set the technical bar for their work, and participate in hiring.
What We’re Looking For
  • Bachelor's degree in Computer Science, Robotics, Electrical Engineering, Statistics, or a related field with 8+ years of relevant experience.
  • 2+ years as a technical lead or engineering manager for an evaluation, validation, or test team.
  • Ownership of an offline or system-level evaluation system for an autonomous vehicle or robotics program, or a comparable large-scale ML model evaluation system in production, with accountability for the metrics and for the release decisions made on them.
  • Experience designing safety or behavior metrics for an autonomous system and defending them to engineering and executive audiences.
  • Hands‑on experience building log replay at scale, including scoring against ground truth in open-loop and closed-loop evaluation.
  • Experience running evaluation pipelines on GPU clusters or comparable large-scale compute on a recurring cadence, with responsibility for reproducibility, throughput, and compute cost.
  • Experience with evaluation dataset design, benchmarking, and regression detection across software or model versions.
  • Working command of applied statistics for system validation, including statistical uncertainty, sampling, rare‑event analysis, and the data volume a given claim requires.
  • Working understanding of autonomous systems end to end, including sensing, perception, localization, planning, and controls.
  • Experience on an autonomous vehicle or robotics platform.
  • Working experience with a robotics middleware and its logging and replay tooling.
  • Experience handling autonomous driving or robotics sensor data across multiple modalities and timestamps, including camera, LiDAR, and radar.
  • Proficiency in Python and the ability to read, navigate, and debug existing C++ codebases.
Why join us

At Atoms, you’ll work on one of the defining challenges of our time - bringing automation into the physical world to drive real, lasting impact. We exist to uncover valuable unknown truths and turn them into progress, which means constantly pushing beyond what’s known and building what doesn’t yet exist. The work is ambitious and often challenging, but it’s grounded in a shared sense of purpose and a team committed to seeing it through together. Our work only matters if it serves others, and we know that meaningful progress depends on the trust of the people we serve and the strength of our team - so we invest in both, creating an environment where you can do your best work and grow.

What Else You Need To Know

This role is based in our San Francisco office location. As a company driven by innovation and continuous change, close collaboration is essential. We’re constantly reimagining our industry, creating new products, and refining our processes, and we do our best work together. That’s why all of our office‑based teams work onsite, five days a week.

The base salary range for this role is $185,000 - $242,000 per year.

Actual compensation will be determined on an individual basis and may vary depending on experience, skills, and qualifications.

Base salary is just one part of your total rewards package. You may also be eligible for equity awards.

Benefits Summary (USA Full-Time Exempt Employees)
  • Medical, Dental, Vision, Disability, and Life Insurance
  • Flexible Spending Account / Health Savings Account Options
  • 401(k)
  • Equity
  • Sick Time, Unlimited Flexible Time Off, and Paid Holidays
  • Paid Parental Leave
  • Pre‑Tax Commuter Benefit Plan
  • Team lunch in our SoMa office every Tuesday and Thursday

Benefits Are Subject To Change At The Company's Discretion.

Atoms accepts applications on an ongoing basis.

Ready to join us as we serve those who serve others?
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Lead - Autonomy Evaluation
Technical Lead - Autonomy Evaluation

ATOMS Careers page • San Francisco (CA)

On-site
USD 185,000 - 242,000
Medical insurance
Dental & Vision
401(k)
+5
Robotics Lead
Robotics Lead

ATOMS Careers page • San Francisco (CA)

On-site
USD 255,000 - 322,000
Medical, Dental, Vision, Disability, &
401(k)
Equity
+3
Robotics Lead
Robotics Lead

Atoms • San Francisco (CA)

On-site
USD 255,000 - 322,000
Equity
401(k)
Paid Holidays
+2
Senior Systems Integration Engineer
Senior Systems Integration Engineer

Atoms • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, Dental, Vision, Disability &
Life Insurance
401(k)
+6
Senior Director, Mining Autonomous Platform (AV Platform Engineering)
Senior Director, Mining Autonomous Platform (AV Platform Engineering)

Cssmerge • San Francisco (CA)

On-site
USD 255,000 - 322,000
Medical, dental, vision insurance
401(k)
Equity
+4
Robotics Systems Engineer
Robotics Systems Engineer

Pronto • San Francisco (CA)

On-site
USD 145,000 - 180,000
Medical insurance
Equity
401(k)
+4
Senior Software Engineer, Simulation
Senior Software Engineer, Simulation

Cssmerge • San Francisco (CA)

On-site
USD 195,000 - 246,000
Equity
401(k)
Parental Leave
+2
Senior Software Engineer, Simulation
Senior Software Engineer, Simulation

ATOMS Careers page • San Francisco (CA)

On-site
USD 195,000 - 246,000
Healthcare
Equity
401(k)
+1
Senior Director, Mining Autonomous Platform (AV Platform Engineering)
Senior Director, Mining Autonomous Platform (AV Platform Engineering)

Atoms • San Francisco (CA)

On-site
USD 255,000 - 322,000
Medical, Dental, Vision
401(k)
Equity
+1
Senior Systems Engineer, Diagnostics and System Behavior Analysis
Senior Systems Engineer, Diagnostics and System Behavior Analysis

Cssmerge • San Francisco (CA)

On-site
USD 176,000 - 242,000
Medical, Dental, Vision insurance
401(k)
Equity
+3