AI Personal-Workflow Evaluator & Rubric Specialist

Mercor

New York (NY)

Remote

USD 34,000 - 83,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is seeking advanced LLM power users to evaluate how AI systems handle complex personal-life tasks using MCP and personal plugins/connectors like Google Drive, Expedia, and Notion. You will design realistic prompts and perform tasks while recording your screen.

You should be US-based, have a rich LLM account with 6+ months history, and be comfortable signing a DocuSign data-share consent. Expect 20+ hours per week, rapid turnarounds, and use of multiple plugins to test multi-step planning

Qualifications

  • US-based only.
  • Strong MCP experience and plug-in/connector usage.
  • Experience using LLM plugins/connectors such as Google Drive, Expedia, Notion, and similar tools, multiple times a week.
  • Heavy personal usage of LLM products.
  • An active, rich LLM account with regular usage and approximately 6+ months of history.
  • Willingness to sign a data-share consent form via DocuSign.
  • Experience using AI for multi-step planning, research, decision-making, or personal workflows.
  • Strong written judgment and attention to detail.
  • Ability to explain what makes an AI output good, bad, incomplete, unsafe, or unrealistic.
  • Experience writing and evaluating against rubrics.
  • Extensive rubric experience is valuable, including 100+ hours on prior rubric projects.

Responsibilities

  • Creating realistic prompts for complex personal-life tasks.
  • Executing tasks and actions while recording your screen (required).
  • Using your personal plugins/connectors while you complete actions.
  • Writing clear explanations of AI successes and failures.
  • Judging whether AI outputs are practical, personalized, and well-reasoned.
  • Identifying where models miss context, overreach, fail to use tools correctly, or produce unrealistic results.
  • Creating and applying detailed rubrics to assess model performance.

Skills

MCP experience
Plugin usage
Contextual reasoning
Multi-step planning
Attention to detail
Strong written judgment

Tools

Google Drive
Expedia
Notion

Job description

Mercor is seeking advanced LLM power users to evaluate how AI systems handle complex personal-life tasks using MCP and personal plugins/connectors like Google Drive, Expedia, and Notion. You will design realistic prompts and perform tasks while recording your screen.

You should be US-based, have a rich LLM account with 6+ months history, and be comfortable signing a DocuSign data-share consent. Expect 20+ hours per week, rapid turnarounds, and use of multiple plugins to test multi-step planning

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Power User & MCP Specialist: Plugin-Driven Tasks
LLM Power User & MCP Specialist: Plugin-Driven Tasks

Mercor • New York (NY)

On-site
USD 60,000 - 90,000
LLM Power User - MCP Specialist
LLM Power User - MCP Specialist

Mercor • New York (NY)

On-site
USD 60,000 - 90,000
AI Workflow Evaluator - AI Trainer
AI Workflow Evaluator - AI Trainer

Obsidian • New York (NY)

On-site
USD 60,000 - 90,000
AI Workflow Evaluator & Personal Task Trainer
AI Workflow Evaluator & Personal Task Trainer

Obsidian • New York (NY)

On-site
USD 60,000 - 90,000
AI Personal Workflow Evaluator
AI Personal Workflow Evaluator

Obsidian • New York (NY)

Remote
USD 34,000 - 55,000
AI Personal Assistant Evaluator
AI Personal Assistant Evaluator

Mercor • San Francisco (CA)

On-site
USD 50,000 - 90,000
AI Personal Assistant Evaluator
AI Personal Assistant Evaluator

Obsidian • San Francisco (CA)

On-site
USD 34,000 - 83,000
AI Personal Assistant Evaluator - Real-World Task Specialist
AI Personal Assistant Evaluator - Real-World Task Specialist

Mercor • San Francisco (CA)

On-site
USD 50,000 - 90,000
LLM Power User - Fully Remote | Upto $200/hr
LLM Power User - Fully Remote | Upto $200/hr

Obsidian • San Francisco (CA)

Remote
USD 34,000 - 83,000
LLM Power User: Real-World Task Evaluator
LLM Power User: Real-World Task Evaluator

Obsidian • San Francisco (CA)

On-site
USD 55,000 - 96,000