Project Cursa - Robot Manipulation Video Annotator (V2)

Annotation Academy

United States

Remote

USD 21,000 - 30,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Welo Data in the United States seeks meticulous annotators to label robot manipulation videos for AI training. You will produce precise, structured, natural-language descriptions across three camera angles, following a detailed style guide.

Responsibilities include segmenting videos, labeling at atomic, subtask, and task levels, and ensuring every moment is described— even when grasps fail or objects slip. Strong written English and sharp observation are essential.

Qualifications

  • Strong written English and sharp observational skills.
  • Ability to follow a structured style guide consistently.
  • Comfort with spatial/mechanical descriptions (left/right, above/below).

Responsibilities

  • Watch robot manipulation videos from three synchronized camera angles.
  • Segment videos and write clear, natural-language descriptions for each segment.
  • Apply labels at atomic, subtask, and task levels for each applicable segment.
  • Ensure every moment is covered by labels at two or more levels, including idle moments.
  • Describe exactly what happens, including failed grasps or slips, with precision.

Skills

Strong written English
Attention to detail
Follow a detailed style guide
Spatial description (left/right, above

Tools

Label Studio

Job description

About the Role

We are looking for detail-oriented annotators to help label robot manipulation videos for AI training purposes. You'll watch short videos of robots performing manipulation tasks (filmed from three synchronized camera angles) and produce precise, structured, natural-language descriptions of the actions taking place. This work directly supports the development of robotics AI models and requires strong written English, sharp observational skills, and the discipline to follow a detailed style guide consistently.

What You'll Do
  1. Watch short robot manipulation videos, each filmed from three synchronized camera views (an overhead view and views from each of the robot's two wrist-mounted cameras).
  2. Break each video into time segments and write clear, natural-language descriptions for each segment.
  3. Apply labels at three levels of detail for each applicable segment:
    • Atomic motion (a few seconds) — a single small movement (e.g., "close fingers around the red handle")
    • Skill / subtask (several seconds to ~20seconds) — a complete, meaningful action (e.g., "pick up the red block by its edge")
    • Task / goal (up to ~1minute) — the overall purpose of a sequence of skills (e.g., "place all blocks in the container")
  4. Ensure every moment of video is covered by a label at two or more of these levels — no gaps, including idle or pause moments.
  5. Accurately describe exactly what happens, including when something doesn't go as planned (a dropped object, a failed grasp, a slipped grip). Precision matters more than making the robot look successful.
  6. Cross-reference all three camera angles: use the overhead view to understand the overall scene and object identity, and the close-up wrist views to confirm exact contact and grasp details.
  7. Follow a detailed style guide covering vocabulary for actions, spatial relationships, object descriptions, and manner of movement, applying it consistently across many episodes.
  8. Participate in periodic calibration sessions to align your labeling with the team and the client's reference examples.
What We're Looking For
Required
  • Strong written English — you'll write dozens of short, precise descriptive sentences per video and need to vary your language rather than repeating the same phrases.
  • Sharp attention to detail — able to distinguish small differences (a successful grasp vs. a fumble, a push vs. a drag, which specific object part is being touched).
  • Comfort following a detailed, structured style guide and applying it consistently, even in ambiguous or edge‑case scenarios.
  • Basic comfort with spatial/mechanical description (left/right, above/below, naming object parts like handles, lids, or edges).
  • Reliable, self-directed work habits — this is often heads‑down work with periodic check‑ins rather than close supervision.
Nice to Have
  • Prior experience with video annotation, data labeling, transcription, or QA work.
  • Familiarity with robotics terminology (grippers, end-effectors, manipulation) — helpful but not necessary, as the style guide is self-contained.
  • Experience with annotation tools such as Label Studio.
Why Join Welo Data?
  • Limitless Flexibility: Project-based opportunities that fit your availability. Choose when and how much you want to contribute—fully remote, with complete autonomy.
  • Limitless Growth: Optional access to AI and Large Language Model workshops designed specifically for professionals like you. No coding required—just your expertise.
  • Limitless Support: Be part of a global contributor community with responsive guidance and support.
  • Real Impact: Apply your expertise in the Legal field to influence the AI systems shaping the future of your industry—while collaborating with data professionals and expanding your skills.
About Welo Data

Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train the world’s most advanced AI systems. We’re building smarter, more human AI with a diverse community in 100+ countries. At Welo Data, Limitless AI. Limitless You. isn't just a slogan—it’s our promise. We build smarter AI through the power of human contribution, offering limitless opportunities for our global community to grow, contribute, and work on their terms.

#L1-CC1

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Robot Manipulation Video Annotator
Remote Robot Manipulation Video Annotator

Annotation Academy • United States

Remote
USD 21,000 - 30,000
Project Perseus | Data Labeling Associate - Tamil Speakers (Human-in-the-Loop AI)
Project Perseus | Data Labeling Associate - Tamil Speakers (Human-in-the-Loop AI)

Welo Global • Sunnyvale (CA)

On-site
USD 39,000 - 55,000
Paid vacation
Paid holidays
Paid sick leave
+8
Project Cursa - Robot Manipulation Video Annotation QC
Project Cursa - Robot Manipulation Video Annotation QC

Annotation Academy • United States

Remote
USD 60,000 - 90,000
English (UK) Data Labeling Associate - NYC based
English (UK) Data Labeling Associate - NYC based

Welo Data • New York (NY)

On-site
USD 63,000 - 77,000
Paid Vacation: 6 days
Medical, Dental, and Vision Insurance
Free Gourmet Food
+1
English (UK) Data Labeling Associate - California based
English (UK) Data Labeling Associate - California based

Welo Data • California (MO)

On-site
USD 68,000 - 72,000
Paid Vacation: 6 days
Medical, Dental, and Vision Insurance
Free Gourmet Food
+2
Video Data Labelling Expert - Remote
Video Data Labelling Expert - Remote

YO AI Labs • Chicago (IL)

Remote
USD 28,000 - 44,000
Video Data Labelling Expert - Remote
Video Data Labelling Expert - Remote

YO AI Labs • Washington

Remote
USD 34,000 - 55,000
Video Data Labelling Expert - Remote
Video Data Labelling Expert - Remote

YO AI Labs • Town of Schroeppel (NY)

Remote
USD 28,000 - 48,000
Video Data Labelling Expert - Remote
Video Data Labelling Expert - Remote

YO AI Labs • Atlanta (GA)

Remote
USD 21,000 - 34,000
Video Data Labelling Expert - Remote
Video Data Labelling Expert - Remote

YO AI Labs • Philadelphia

Remote
USD 34,000 - 55,000