Language Model Evaluator - Fully Remote | Upto $20/hr Part-time

Mercor

United States

Remote

USD 20,664 - 27,552

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is seeking a Generalist proficient in English and Urdu to conduct evaluations for AI models from a remote location. The successful candidate will engage in high-quality data generation, fact-checking, and assessment of AI responses. A Bachelor's degree and fluency in Urdu are essential.

This position demands attention to detail and significant experience with large language models. The compensation ranges from $15 to $20 per hour.

Qualifications

  • Native speaker in Urdu with excellent writing skills in English.
  • Significant experience using large language models (LLMs).
  • Strong attention to detail and background in analytical thinking.

Responsibilities

  • Conduct fact-checking using trusted public sources and external tools.
  • Generate high-quality human evaluation data and assess responses.
  • Ensure model responses align with expected behavior and guidelines.

Skills

Native speaker in Urdu
Excellent writing skills in English
Strong attention to detail
Significant experience using large language models (LLMs)
Background in structured analytical thinking

Education

Bachelor's degree

Job description

About The Job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.

Position

Generalist - English & Urdu

Type

Contract

Compensation

$15–$20/hour

Location

Remote

Role Responsibilities
  • Conduct fact-checking using trusted public sources and external tools.
  • Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and factual inaccuracies.
  • Assess reasoning quality, clarity, tone, and completeness of responses.
  • Ensure model responses align with expected conversational behavior and system guidelines.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.
Qualifications
Must-Have
  • Bachelor's degree
  • Native speaker in Urdu
  • Significant experience using large language models (LLMs)
  • Excellent writing skills in English
  • Strong attention to detail
  • Background or experience in domains requiring structured analytical thinking
Preferred
  • Experience with RLHF, model evaluation, or data annotation work
  • Experience writing or editing high-quality written content
  • Experience comparing multiple outputs and making fine-grained qualitative judgments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Bilingual LLM Evaluator (Urdu-English)
Remote Bilingual LLM Evaluator (Urdu-English)

Mercor • United States

Remote
Urdu Language Evaluator & AI Prompt Architect (Remote)
Urdu Language Evaluator & AI Prompt Architect (Remote)

Welo Global • Vero Beach (FL)

On-site
Urdu LLM Evaluator - Global Contract Role
Urdu LLM Evaluator - Global Contract Role

Obsidian • San Francisco (CA)

Remote
USD 40,000 - 60,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
Remote English Linguistic Expert for AI Training
Remote English Linguistic Expert for AI Training

Mercor • San Jose (CA)

On-site
USD 60,000 - 100,000
Fully remote
Flexible hours
Asynchronous work
+1
Generative AI Evaluator | $30/hr Remote
Generative AI Evaluator | $30/hr Remote

Crossing Hurdles • United States

Remote
English Language Expert
English Language Expert

SME Careers • Wichita (KS)

On-site
USD 34,000 - 55,000
English Language Expert
English Language Expert

SME Careers • Portland (ME)

On-site
USD 34,000 - 55,000
English Language Expert
English Language Expert

SME Careers • Bloomington (IN)

On-site
USD 28,000 - 50,000
C++ Systems Engineer – AI Model Evaluation & Code Review | Remote
C++ Systems Engineer – AI Model Evaluation & Code Review | Remote

Crossing Hurdles • United States

On-site