Remote LLM Evaluator & Prompt Engineer (Contract)

Benture

San Francisco (CA)

Remote

USD 827,000 - 3,306,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Turing is seeking detail-oriented LLM Annotators for a remote, short-term contract role. You will evaluate and improve cutting-edge language models, design prompts, and validate outputs against source data.

Ideal candidates hold a Master's degree and have strong analytical and communication skills. The position offers flexible hours (20, 30, or 40 per week) with overlap for collaboration, and no medical or paid leave benefits as a contractor role.

Qualifications

  • Master's degree or higher in any discipline.
  • Minimum 3 years of professional, research, or teaching experience.
  • Excellent written English communication skills.
  • Exceptional attention to detail and ability to validate information against source data.

Responsibilities

  • Design challenging prompts that evaluate an LLM's ability to retrieve, analyze, and reason over structured data.
  • Assess AI-generated responses for factual accuracy, logical reasoning, and completeness.
  • Identify model failures, hallucinations, inconsistencies, and reasoning gaps.
  • Validate model outputs using provided datasets and supporting evidence.
  • Document findings with clear, evidence-based explanations.
  • Adhere consistently to annotation guidelines and maintain high-quality standards.

Skills

Analytical thinking
Attention to detail
Written English

Education

Master's degree or higher

Tools

CSV
Excel
Databases

Job description

Turing is seeking detail-oriented LLM Annotators for a remote, short-term contract role. You will evaluate and improve cutting-edge language models, design prompts, and validate outputs against source data.

Ideal candidates hold a Master's degree and have strong analytical and communication skills. The position offers flexible hours (20, 30, or 40 per week) with overlap for collaboration, and no medical or paid leave benefits as a contractor role.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Annotator - Master's Degree · Turing Turing · Varies · remote · 2w ago Varies 2w ago
LLM Annotator - Master's Degree · Turing Turing · Varies · remote · 2w ago Varies 2w ago

Benture • San Francisco (CA)

Remote
USD 827,000 - 3,306,000
Remote English NLP Editor & Linguistic Annotator (Contract)
Remote English NLP Editor & Linguistic Annotator (Contract)

Mercor • San Francisco (CA)

Hybrid
GBP 46,000 - 56,000
Vietnamese LLM Prompt & Evaluation Specialist (Remote)
Vietnamese LLM Prompt & Evaluation Specialist (Remote)

Benture • San Francisco (CA)

Remote
USD 34,440,000 - 82,656,000
Fully remote
Flexible hours
Competitive compensation
Remote LLM Evaluation Specialist - Part-Time
Remote LLM Evaluation Specialist - Part-Time

OpenTrain AI, Inc. • Northern (KY)

Hybrid
USD 45,000 - 65,000
LLM Annotator – Vietnamese · Turing Turing · Varies · remote · 1w ago Varies 1w ago
LLM Annotator – Vietnamese · Turing Turing · Varies · remote · 1w ago Varies 1w ago

Benture • San Francisco (CA)

Remote
USD 34,440,000 - 82,656,000
Fully remote
Flexible hours
Competitive compensation
Multilingual AI Language Analyst (Remote Contract)
Multilingual AI Language Analyst (Remote Contract)

Turing • United States

Remote
USD 34,000 - 69,000
Competitive pay
Flexible hours
Remote work environment
+2
Remote AI Analyst for LLM Training & Analysis
Remote AI Analyst for LLM Training & Analysis

turing • United States

Remote
USD 39,000 - 58,000
Flexible working hours
Remote work environment
Work on cutting-edge AI projects
LLM Tuning Analyst (Remote, 40h/wk)
LLM Tuning Analyst (Remote, 40h/wk)

turing • United States

Remote
USD 34,000 - 55,000
Remote work
Flexible hours
Senior Remote Python Engineer - AI/LLM Evaluation Lead
Senior Remote Python Engineer - AI/LLM Evaluation Lead

turing • United States

Remote
USD 69,000 - 124,000
Fully remote
Cutting-edge AI projects
Remote LLM AI Quality Analyst — Personalization
Remote LLM AI Quality Analyst — Personalization

Donato Technologies Inc • Moraga (CA)

Remote
USD 17,000 - 25,000
Remote work