Research Evaluation Lead

Socket.dev

Mountain View (CA)

On-site

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Rhoda AI is seeking a researcher to own end-to-end eval workflows for cutting-edge robot models. You will translate research questions into clear protocols, run pilots, and drive QA, ensuring consistent results across shifts and stations.

The role emphasizes hands-on coding, rigorous experimentation, and the ability to distinguish model failures from hardware, data, or setup issues. Collaboration with multiple teams is essential.

Qualifications

  • Familiarity with robotics or ML experimentation.
  • Background in computer science or hands-on coding experience.
  • Detail-oriented and able to interpret experimental intent.
  • Demonstrated ownership and execution capability.
  • Experience with robotics testing, data collection, lab operations, or QA preferred.

Responsibilities

  • Translate research intent into eval protocols and success criteria.
  • Own end-to-end evaluation workflow from model handoff to QA and results.
  • Train and manage eval pilots across shifts and stations.
  • Distinguish model failures from hardware, setup, or data issues.
  • Maintain eval setups, resets, randomization, and experiment traceability.
  • Track quality, throughput, and bottlenecks; improve the eval process.
  • Collaborate with Research, Robot Data, and Eval Platform teams.

Skills

Robotics / ML experimentation
Coding
Detail-oriented
Ownership
Robotics testing / QA

Education

Computer science background

Job description

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

Mission

Own end-to-end robot model evaluation for Research. Turn research questions into consistent, high-quality, repeatable evals that enable fast iteration and trusted results.

Responsibilities
  • Translate research intent into clear eval protocols, trial plans, and success criteria.
  • Own execution end-to-end: model handoff → station readiness → pilot execution → QA → results.
  • Train and manage eval pilots; ensure consistency across people, shifts, and stations.
  • Distinguish model failures from hardware, setup, operator, or data-quality issues.
  • Maintain eval setups, resets, randomization, metadata, and experiment traceability.
  • Track quality, throughput, and bottlenecks; continuously improve the eval process.
  • Partner closely with Research, Robot Data, and Eval Platform teams.
What we’re looking for
  • Some understanding of robotics / ML experimentation.
  • Computer science background or hands-on experience with coding
  • Rigorous, detail-oriented, and able to understand the intent behind an experiment, not just execute instructions.
  • Strong hands-on execution and ownership.
  • Experience in robotics testing, data collection, lab operations, or QA preferred.
Success looks like

A researcher can hand off a model and research question and receive a trusted, standardized eval result with sufficient trials and QA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Evaluation Lead
Research Evaluation Lead

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 210,000
Robotics Research Evaluation Lead
Robotics Research Evaluation Lead

Socket.dev • Mountain View (CA)

On-site
USD 120,000 - 180,000
Robotics Research Evaluation Lead
Robotics Research Evaluation Lead

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 210,000
Research Member of Technical Staff- Robot Learning Systems & Reliability
Research Member of Technical Staff- Robot Learning Systems & Reliability

Rhoda AI • Mountain View (CA)

On-site
USD 190,000 - 270,000
Research Member of Technical Staff- Robot Learning Systems & Reliability
Research Member of Technical Staff- Robot Learning Systems & Reliability

Socket.dev • Mountain View (CA)

On-site
USD 180,000 - 320,000
Research Member of Technical Staff- Robot Learning Systems & Reliability
Research Member of Technical Staff- Robot Learning Systems & Reliability

Rhoda • Mountain View (CA)

On-site
USD 180,000 - 240,000
Research Member of Technical Staff- Post-training & Robot Learning
Research Member of Technical Staff- Post-training & Robot Learning

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
Senior Data Project Manager
Senior Data Project Manager

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Research, Post-Training Evals
Research, Post-Training Evals

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Research, Post-Training Evals
Research, Post-Training Evals

Mosaic.tech • San Francisco (CA)

On-site
USD 140,000 - 190,000