Senior Annotation and Data Pipeline Manager

Genesis AI

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Genesis AI is looking for an experienced professional to scale and manage their data engine and annotation pipelines in San Francisco, California. This role involves overseeing the process from raw data to training-ready datasets and requires a deep understanding of data and machine learning operations.

The ideal candidate will have 4+ years leading data pipelines, strong Python skills, and the ability to automate processes effectively. Join a pioneering AI lab that emphasizes hands-on technical leadership in a fast-paced environment.

Qualifications

  • 4+ years in data or ML pipelines with leadership experience.
  • Strong Python skills (Pandas, NumPy, PyTorch) and SQL.
  • Understand ML concepts like training versus test, precision, and recall.
  • Experience in hands-on technical leadership.

Responsibilities

  • Run the data engine and ensure training-ready datasets.
  • Design the ontology with the model team.
  • Automate trajectory annotation and data synthesis.
  • Stand up and scale labeling with a delivery schedule.
  • Transform evaluation failures into targeted collection jobs.
  • Track metrics like inter-annotator agreement and label error rate.

Job description

The role

We have built a frontier model and put Eno in front of the world, fast. Behind that is a data engine: the machine that turns a raw human demonstration into data the model is measurably better for. This role owns that engine.

A worn glove and a camera produce a raw demonstration, not training data. You will build the pipeline and the annotation operation that turn raw demonstrations into clean, labeled, training-ready data, and make it scale with automation rather than headcount. You will own the datasets, what gets annotated, and the ontology, how it gets labeled, bring vision-language models to bear on trajectory labeling and language grounding, and close the loop so the engine keeps making the model better. This role serves the whole operation, our own floors and our partner-funded collection.

What you'll do
  • Run the data engine. Own the loop from raw trajectory and video to training-ready datasets, with validation steps that guarantee clean, correctly labeled data.

  • Own datasets and ontology. Decide what gets annotated and how, designing the ontology with the model team for its training implications.

  • Automate with models. Use vision-language models for automated trajectory annotation, language grounding, and data synthesis, so the pipeline scales without linear headcount, while holding the quality bar.

  • Run the annotation operation. Stand up and scale labeling, internal and vendor, against a clear quality bar and a delivery schedule the model team can plan around.

  • Close the loop. Turn real-robot eval failures into targeted collection and annotation jobs, and prove the new data improves the model.

  • Own the metrics. Track inter-annotator agreement, label error rate, and throughput per annotator-hour, and drive them the right way.

What we're looking for
  • You have scaled an annotation or data pipeline at a serious operation. Four or more years in data or ML pipelines, including time leading the work. At a frontier AI lab or a top data operation, you have taken raw robot or embodied data to training-ready at volume and you know exactly where it breaks. The people who have done this are a small group. If you are one, we want to talk.

  • You can build, not just manage. Strong Python (Pandas, NumPy, PyTorch) and SQL. You write the automation that shrinks the pipeline.

  • ML literacy. You understand training versus test, precision and recall, and overfitting well enough to design an ontology that helps the model, not just labels data.

  • Hands-on technical leadership. You can run a labeling operation and stay a hands-on contributor at the same time.

  • Comfortable with ambiguity and speed. You move fast in a research-paced environment and bring order to it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Annotation Lead
Data Annotation Lead

Physical Intelligence • San Francisco (CA)

On-site
USD 120,000 - 160,000
Annotation Operations Manager
Annotation Operations Manager

Dyna Robotics • Redwood City (CA)

On-site
USD 110,000 - 170,000
Annotation Operations Manager Redwood City, CA Fulltime
Annotation Operations Manager Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 90,000 - 130,000
Data Annotation Lead
Data Annotation Lead

Kindredventures • United States

On-site
USD 90,000 - 140,000
Senior Data Operations Manager
Senior Data Operations Manager

Genesis AI • United States

On-site
USD 90,000 - 130,000
Data Operations Manager
Data Operations Manager

Genesis AI • San Francisco (CA)

On-site
USD 90,000 - 130,000
Staff ML Engineer, Agent Training & Environments
Staff ML Engineer, Agent Training & Environments

EngineersOfAI • San Francisco (CA)

On-site
USD 170,000 - 260,000
AI Data Strategist
AI Data Strategist

Dyna Robotics • Redwood City (CA)

On-site
USD 100,000 - 140,000
Applied Research Engineer
Applied Research Engineer

HRB • San Francisco (CA)

On-site
USD 180,000 - 280,000
Technical Program Manager, Data Engine
Technical Program Manager, Data Engine

Sunday • Redwood City (CA)

On-site
USD 90,000 - 120,000