Remote AI QA Trainer: LLM Evaluation & Reliability

Agency

United States

Remote

USD 8,300 - 90,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Agency seeks AI QA trainers to rigorously evaluate large language models across multilingual and domain-specific contexts. You will design tests, run regression suites, and document failure modes to raise the bar for model reasoning, reliability, and safety.

You’ll work on adversarial red-teaming, automation with Python/SQL, and dashboards to track quality deltas. A strong background in ML/AI QA and tooling signals fit.

Qualifications

  • A bachelor’s, master’s, or PhD in CS/DS/CL/ statistics is ideal.
  • Shipped QA for ML/AI systems and safety/red-team experience.
  • Experience with test automation frameworks (e.g., PyTest).
  • Hands-on with LLM eval tooling (OpenAI Evals, RAG evaluators, W&B).

Responsibilities

  • Converse with the model on real-world scenarios and evaluation prompts, verify factual accuracy and logical soundness.
  • Design and run test plans and regression suites, build clear rubrics and pass/fail criteria.
  • Capture reproducible error traces with root-cause hypotheses, and suggest improvements to prompt engineering, guardrails, and evaluation metrics.
  • Partner on adversarial red-teaming, automation (Python/SQL), and dashboarding to track quality deltas over time.

Skills

Model evaluation
LLM safety
Prompt robustness
Data quality assurance
Multilingual testing
Grounding verification
Compliance checks

Education

CS/ML degree
PhD ideal

Tools

Python
SQL

Job description

Agency seeks AI QA trainers to rigorously evaluate large language models across multilingual and domain-specific contexts. You will design tests, run regression suites, and document failure modes to raise the bar for model reasoning, reliability, and safety.

You’ll work on adversarial red-teaming, automation with Python/SQL, and dashboards to track quality deltas. A strong background in ML/AI QA and tooling signals fit.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

On-site
USD 8,265 - 89,544
Secure computer and high-speed internet required
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Agency • United States

Remote
USD 8,300 - 90,000
Remote AI QA Trainer & Evaluation Specialist
Remote AI QA Trainer & Evaluation Specialist

Meridial • United States

Remote
USD 8,265 - 89,544
Secure computer and high-speed internet required
Remote AI QA Trainer & Evaluation Specialist
Remote AI QA Trainer & Evaluation Specialist

Meridial • United States

Remote
USD 8,265 - 89,544
Secure computer and high-speed internet required
Remote QA Engineer for AI Evaluation & Testing
Remote QA Engineer for AI Evaluation & Testing

24-MAG • United States

Remote
USD 124,000 - 241,000
AI Training & Evaluation Specialist
AI Training & Evaluation Specialist

MCI • United States

On-site
USD 70,000 - 90,000
Remote Linguistic QA Reviewer (Contract)
Remote Linguistic QA Reviewer (Contract)

AI Trainer Jobs • United States

Remote
USD 110,000 - 131,000
LLM QA Engineer: AI Testing & Evaluation
LLM QA Engineer: AI Testing & Evaluation

Codefeast • United States

On-site
USD 90,000 - 140,000
Senior AI Engineer (LLM Training & RLHF) - Remote
Senior AI Engineer (LLM Training & RLHF) - Remote

Prolific • Virginia Beach (VA)

On-site
USD 100,000 - 140,000
Competitive pay rates
Flexible hours
Ability to work from home
Remote AI Language Evaluation Specialist
Remote AI Language Evaluation Specialist

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000