- Nimble Brain turns real-world operational data into the training sets, evaluations, and feedback loops that make our robots smarter every day. The fuel for that engine is human demonstration data — skilled operators across our sites performing and recording the tasks our superhumanoids learn from — and the quality of every episode we collect sets the ceiling on every model we train
- We’re looking for a Manager, Data Quality & Annotation to become the first dedicated owner of that quality bar. Own the rubric, not just run it: you’ll write the acceptance criteria that define what a “good” episode is for every task family, build the audit workflows that enforce them as we scale from ~15 to 100+ operators across four sites this year, and stand up the annotation engine — including a remote annotation team turning around episode review overnight — that keeps labeled, trusted data flowing to research on schedule
- This is a hands‑on, metrics‑driven, build-from-scratch role. You’ll be based at our San Francisco HQ and spend heavy time at our collection sites, especially during operator ramps. Success looks like an episode acceptance rate that holds steady while the operator base grows 10x — with rubrics, audits, and dashboards that run like clockwork instead of heroics
- Own episode acceptance criteria for every task family — translate research and engineering data needs into operational, auditable standards, and version them as model training needs evolve
- Take over the annotation and episode‑review queue in your first weeks, then build the team and workflows that scale it far beyond yourself
- Stand up and manage a remote annotation team (likely Philippines‑based; direct hires or vendor/BPO) with an overnight turnaround SLA — episodes collected today are reviewed and scored before the next shift starts
- Design and run a sampling‑based QA audit program with explicit coverage targets, including inter‑rater reliability checks that keep annotators and auditors calibrated
- Run the operator quality feedback loop: operator‑level quality scorecards delivered to site supervisors within 24 hours. You own the standard and the signal; supervisors own the coaching and people decisions
- Instrument your function: define the quality metrics (episode acceptance rate, audit coverage, feedback latency, operator quality distribution, annotation throughput) and build the operational dashboards your team runs on, partnering with our analytics function, which independently owns org‑wide reporting
- Own the certification bar for new operator onboarding — no one collects production data without meeting it — while site teams run the day‑to‑day training reps
- Drive a standing weekly loop with research and engineering on failure modes, task‑spec drift, and what “good” needs to mean next
A track record of building quality standards from scratch — rubrics, SOPs, QC workflows — not just executing against existing ones. Be ready to walk us through one you builtMeticulous judgment on edge cases, paired with the pragmatism to ship a v1 rubric this week instead of a perfect one next quarterData fluency: able to build and own your own reporting and pressure‑test the numbers — SQL, BI tools, or AI‑assisted, we don’t care how — rather than waiting on someone elseExperience managing annotation or review teams, including remote/offshore or vendor/BPO teams, and driving their performance against SLAsHigh agency and comfort with ambiguity in a fast‑paced, high‑growth environment — the org will triple around you this yearAble to work in person out of our San Francisco HQ, with regular time at our collection sites3+ years in data operations, annotation/labeling operations, or data collection QA for ML systems — robotics, autonomy, or teleoperation data strongly preferredExperience delivering direct, frequent quality feedback to operators or annotators — including the hard conversations when someone isn’t meeting the barAlignment with Nimble’s values: relentlessly resourceful, humble, dependable, and committed to legendary impactFamiliarity with imitation learning / robot learning data — what makes a demonstration usable for model trainingExperience selecting and managing offshore annotation vendors, or standing up direct offshore teams (EOR, timezone‑shifted workflows)Experience scaling annotation or rater teams: pods, shift leads, inter‑rater reliability programsMulti‑site operations experience