LLM Power User: Real-World Task Evaluator

Obsidian

San Francisco (CA)

On-site

USD 55,000 - 96,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Obsidian in San Francisco is seeking advanced LLM power users to evaluate how well AI systems handle personalized, real-world life tasks. You will judge usefulness, personalization, realism, and safety across diverse domains such as food, health, productivity, careers, and learning.

Expect a commitment of 20-40 hours per week, with flexible scheduling to fit your personal workflows. Your detailed feedback will directly inform improvements to AI assistants, making them more trustworthy and

Qualifications

  • Heavy personal usage of LLM products.
  • Experience using AI for multi-step tasks, planning, research, decision-making, or personal workflows.
  • Familiarity with tools such as ChatGPT, Claude, Gemini, Perplexity, Cursor, Windsurf, Codex, or other AI agents.
  • Ability to explain what makes an AI output good, bad, incomplete, unsafe, or unrealistic.
  • Strong written judgment and attention to detail.

Responsibilities

  • Evaluate how well AI systems handle personalized, real-world tasks.
  • Provide clear judgments on usefulness, personalization, realism, and safety of AI outputs.
  • Document findings to help improve AI assistants for real-world workflows.

Skills

Heavy personal usage of LLMs
Strategic planning
Strong written judgment
Context personalization awareness

Tools

ChatGPT
Claude
Gemini
Perplexity
Cursor
Codex

Job description

Obsidian in San Francisco is seeking advanced LLM power users to evaluate how well AI systems handle personalized, real-world life tasks. You will judge usefulness, personalization, realism, and safety across diverse domains such as food, health, productivity, careers, and learning.

Expect a commitment of 20-40 hours per week, with flexible scheduling to fit your personal workflows. Your detailed feedback will directly inform improvements to AI assistants, making them more trustworthy and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Power User & Real-World AI Evaluator
LLM Power User & Real-World AI Evaluator

Obsidian • San Francisco (CA)

Remote
USD 34,000 - 83,000
LLM Power User - Fully Remote | Upto $200/hr
LLM Power User - Fully Remote | Upto $200/hr

Obsidian • San Francisco (CA)

Remote
USD 34,000 - 83,000
LLM Power User - AI Trainer
LLM Power User - AI Trainer

Obsidian • San Francisco (CA)

On-site
USD 55,000 - 96,000
AI Personal Assistant Evaluator
AI Personal Assistant Evaluator

Mercor • San Francisco (CA)

On-site
USD 50,000 - 90,000
AI Personal Assistant Evaluator
AI Personal Assistant Evaluator

Obsidian • San Francisco (CA)

On-site
USD 34,000 - 83,000
AI Personal Assistant Quality Evaluator
AI Personal Assistant Quality Evaluator

Obsidian • San Francisco (CA)

On-site
USD 34,000 - 83,000
AI Personal Assistant Evaluator - Real-World Task Specialist
AI Personal Assistant Evaluator - Real-World Task Specialist

Mercor • San Francisco (CA)

On-site
USD 50,000 - 90,000
Senior Personal AI Assistant Evaluator
Senior Personal AI Assistant Evaluator

Dorado • United States

Remote
USD 34,000 - 55,000
LLM Power User - MCP Specialist
LLM Power User - MCP Specialist

Mercor • New York (NY)

On-site
USD 60,000 - 90,000
AI Personal-Workflow Evaluator & Rubric Specialist
AI Personal-Workflow Evaluator & Rubric Specialist

Mercor • New York (NY)

Remote
USD 34,000 - 83,000