AI Workflow Evaluator - AI Trainer

Obsidian

New York (NY)

On-site

USD 60,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking advanced LLM power users to evaluate AI systems on real-world personal-life workflows. You will use MCP and plugins/connectors like Google Drive, Notion, and Expedia to perform multi-step tasks such as travel planning, health research, home services, and career planning.

The role requires documenting decisions, recording your screen, and articulating why AI outputs are strong, weak, or unrealistic.

Qualifications

  • US-based only.
  • Strong MCP experience and plug-in/connector usage.
  • Experience using LLM plugins/connectors such as Google Drive, Expedia, Notion, and similar tools, multiple times a week.
  • Heavy personal usage of LLM products.
  • An active, rich LLM account with regular usage and approximately 6+ months of history.
  • Willingness to sign a data-share consent form via DocuSign.
  • Experience using AI for multi-step planning, research, decision-making, or personal workflows.
  • Strong written judgment and attention to detail.
  • Ability to explain what makes an AI output good, bad, incomplete, unsafe, or unrealistic.
  • Experience writing and evaluating against rubrics.

Responsibilities

  • Evaluate AI systems on complex personal workflows and real-world tasks.
  • Create realistic prompts for multi-step life tasks and record your screen during execution.
  • Use personal plugins/connectors to complete actions and document outcomes.
  • Write clear explanations of AI successes and failures and assess practicality and personalization.
  • Identify gaps where models miss context or misapply tools and develop rubrics for evaluation.

Skills

MCP experience
Plugin usage
LLM productivity

Tools

Google Drive
Notion
Expedia

Job description

About the Opportunity

A leading AI research organization is seeking advanced LLM power users with strong experience using MCP and (more importantly) plugins/connectors for real-world personal life tasks.

This project focuses on evaluating how well AI systems handle personalized, multi-step life tasks that require context, judgment, planning, and use of connected tools such as Google Drive, Expedia, Notion, and other plugins/connectors.

This role is ideal for people who use AI heavily in their personal lives and can do a better job replicating what AI could do if it weren't available

What You'll Do

You will help evaluate AI systems on complex personal workflows, including tasks across:

  • Personal health

  • Travel

  • Activity planning, including food and dining

  • Services, such as home repair

  • Career search

  • Other life organization workflows

Responsibilities may include:

  • Creating realistic prompts for complex personal-life tasks

  • Executing tasks and actions while recording your screen (required)

  • Using your personal plugins/connectors while you complete actions

  • Writing clear explanations of AI successes and failures

  • Judging whether AI outputs are practical, personalized, and well-reasoned

  • Identifying where models miss context, overreach, fail to use tools correctly, or produce unrealistic results

  • Creating and applying detailed rubrics to assess model performance

Who We're Looking For

Strong candidates will have:

  • US-based only

  • Strong MCP experience and plug in / connector usage

  • Experience using LLM plugins/connectors such as Google Drive, Expedia, Notion, and similar tools, multiple times a week

  • Heavy personal usage of LLM products

  • An active, rich LLM account with regular usage and approximately 6+ months of history

  • Willingness to sign a data-share consent form via DocuSign

  • Experience using AI for multi-step planning, research, decision-making, or personal workflows

  • Strong written judgment and attention to detail

  • Ability to explain what makes an AI output good, bad, incomplete, unsafe, or unrealistic

  • Experience writing and evaluating against rubrics

Extensive rubric experience is especially valuable, including 100+ hours on prior rubric projects involving rubric design, evaluation, and quality assessment.

Ideal Candidate Profile

The strongest candidates are LLM power users who are already using plug-in tools in their personal lives for high-context tasks such as trip planning, health research, home services, food and dining decisions, career planning, personal organization, or similar workflows.

Candidates who want to be more competitive for this and future opportunities are encouraged to proactively spend time learning MCP and using LLM plugins/connectors before applying.

Why This Work Matters

LLMs are quickly becoming personal assistants for everyday decisions, but truly useful AI needs to do more than produce generic advice. It needs to understand context, preferences, constraints, tradeoffs, and what success looks like in real life.

Your evaluations will help improve how AI systems support people with practical, high-context tasks across food, health, travel, productivity, careers, and life organization. This work directly contributes to making AI assistants more personalized, trustworthy, and useful for real-world personal workflows.

Engagement Details
  • Expected commitment: 20+ hours/week

  • Ramp-up: 1-2 days required

  • Turnaround expectation: Ability to complete tasks within 24 hours

  • Equipment: Desktop or laptop required; Chromebooks are not supported

  • Experts added to the project will begin in a trial period to assess project fit, quality, and consistency before being considered for ongoing tasking.

  • Please note: This project is still in its early stages, so there may be an initial delay before tasking begins.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Power User - MCP Specialist
LLM Power User - MCP Specialist

Mercor • New York (NY)

On-site
USD 60,000 - 90,000
AI Personal-Workflow Evaluator & Rubric Specialist
AI Personal-Workflow Evaluator & Rubric Specialist

Mercor • New York (NY)

Remote
USD 34,000 - 83,000
AI Workflow Evaluator - Fully Remote | Upto $190/hr
AI Workflow Evaluator - Fully Remote | Upto $190/hr

United States Digital Space LLC • United States

Remote
USD 69,000 - 262,000
AI Workflow Evaluator & Personal Task Trainer
AI Workflow Evaluator & Personal Task Trainer

Obsidian • New York (NY)

On-site
USD 60,000 - 90,000
AI Personal Workflow Evaluator
AI Personal Workflow Evaluator

Obsidian • New York (NY)

Remote
USD 34,000 - 55,000
AI Workflow Evaluator & Plugin Expert (Remote)
AI Workflow Evaluator & Plugin Expert (Remote)

United States Digital Space LLC • United States

Remote
USD 69,000 - 262,000
LLM Power User & MCP Specialist: Plugin-Driven Tasks
LLM Power User & MCP Specialist: Plugin-Driven Tasks

Mercor • New York (NY)

On-site
USD 60,000 - 90,000
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

Hybrid
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
Personalized AI Response Evaluator
Personalized AI Response Evaluator

OpenTrain AI • Northern (KY)

Hybrid
USD 21,000 - 34,000
Senior AI Engineer
Senior AI Engineer

Apt • Dallas (TX)

On-site
USD 120,000 - 190,000