AI QA Trainer - LLM Evaluation - Freelance Project

Agency

United States

Remote

USD 8,300 - 90,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Agency seeks AI QA trainers to rigorously evaluate large language models across multilingual and domain-specific contexts. You will design tests, run regression suites, and document failure modes to raise the bar for model reasoning, reliability, and safety.

You’ll work on adversarial red-teaming, automation with Python/SQL, and dashboards to track quality deltas. A strong background in ML/AI QA and tooling signals fit.

Qualifications

  • A bachelor’s, master’s, or PhD in CS/DS/CL/ statistics is ideal.
  • Shipped QA for ML/AI systems and safety/red-team experience.
  • Experience with test automation frameworks (e.g., PyTest).
  • Hands-on with LLM eval tooling (OpenAI Evals, RAG evaluators, W&B).

Responsibilities

  • Converse with the model on real-world scenarios and evaluation prompts, verify factual accuracy and logical soundness.
  • Design and run test plans and regression suites, build clear rubrics and pass/fail criteria.
  • Capture reproducible error traces with root-cause hypotheses, and suggest improvements to prompt engineering, guardrails, and evaluation metrics.
  • Partner on adversarial red-teaming, automation (Python/SQL), and dashboarding to track quality deltas over time.

Skills

Model evaluation
LLM safety
Prompt robustness
Data quality assurance
Multilingual testing
Grounding verification
Compliance checks

Education

CS/ML degree
PhD ideal

Tools

Python
SQL

Job description

Are you an AI QA expert eager to shape the future of AI? Large-scale language models are evolving from clever chatbots into enterprise-grade platforms. With rigorous evaluation data, tomorrow’s AI can democratize world-class education, keep pace with cutting-edge research, and streamline workflows for teams everywhere. That quality begins with you—we need your expertise to harden model reasoning and reliability.

We’re looking for AI QA trainers who live and breathe

  • model evaluation
  • LLM safety
  • prompt robustness
  • data quality assurance
  • multilingual and domain-specific testing
  • grounding verification
  • compliance/readiness checks

You’ll challenge advanced language models on tasks like

  • hallucination detection
  • factual consistency
  • prompt-injection and jailbreak resistance
  • bias/fairness audits
  • chain-of-reasoning reliability
  • tool‑use correctness
  • retrieval‑augmentation fidelity
  • end‑to‑end workflow validation

—documenting every failure mode so we can raise the bar.

On a typical day, you will

  • converse with the model on real-world scenarios and evaluation prompts, verify factual accuracy and logical soundness
  • design and run test plans and regression suites, build clear rubrics and pass/fail criteria
  • capture reproducible error traces with root‑cause hypotheses, and suggest improvements to prompt engineering, guardrails, and evaluation metrics (e.g., precision/recall, faithfulness, toxicity, and latency SLOs)
  • partner on adversarial red‑teaming, automation (Python/SQL), and dashboarding to track quality deltas over time

A bachelor’s, master’s, or PhD in computer science, data science, computational linguistics, statistics, or a related field is ideal; shipped QA for ML/AI systems, safety/red‑team experience, test automation frameworks (e.g., PyTest), and hands‑on work with LLM eval tooling (e.g., OpenAI Evals, RAG evaluators, W&B) signal fit.

Skills that stand out include

  • evaluation rubric design
  • adversarial testing/red‑teaming
  • regression testing at scale
  • bias/fairness auditing
  • grounding verification
  • prompt and system‑prompt engineering
  • test automation (Python/SQL)
  • high‑signal bug reporting

Clear, metacognitive communication—“showing your work”—is essential.

Ready to turn your QA expertise into the quality backbone for tomorrow’s AI?

We offer a pay range of $6-to- $65 per hour, with the exact rate determined after evaluating your experience, expertise, and geographic location. Final offer amounts may vary from the pay range listed above.

As a contractor you’ll supply a secure computer and high-speed internet; company-sponsored benefits such as health insurance and PTO do not apply.

Employment type: ContractWorkplace type: RemoteSeniority level: Mid-Senior Level

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

On-site
USD 8,265 - 89,544
Secure computer and high-speed internet required
Remote AI QA Trainer: LLM Evaluation & Reliability
Remote AI QA Trainer: LLM Evaluation & Reliability

Agency • United States

Remote
USD 8,300 - 90,000
Remote | QA Engineer — $90–$175/hour
Remote | QA Engineer — $90–$175/hour

engineeringjobs.net, Inc. • United States

Remote
USD 124,000 - 241,000
English Quality Assurance Lead (QAL)
English Quality Assurance Lead (QAL)

AI Trainer Jobs • United States

Remote
USD 29,000 - 48,000
Remote contractor role
Remote AI QA Trainer & Evaluation Specialist
Remote AI QA Trainer & Evaluation Specialist

Meridial • United States

Remote
USD 8,265 - 89,544
Secure computer and high-speed internet required
Remote AI QA Trainer & Evaluation Specialist
Remote AI QA Trainer & Evaluation Specialist

Meridial • United States

Remote
USD 8,265 - 89,544
Secure computer and high-speed internet required
Data Scientist Team Lead
Data Scientist Team Lead

AI Trainer Jobs • United States

Remote
USD 257,887,000 - 372,503,000
AI QA Engineer
AI QA Engineer

Cavendish Professionals • Town of Italy (NY)

On-site
USD 95,000 - 120,000
100% Remote – QA Automation OR Data Scientist with AI Exp.
100% Remote – QA Automation OR Data Scientist with AI Exp.

SDH Systems • United States

Remote
USD 120,000 - 155,000
QA Engineer | Remote | $30-$60/hr
QA Engineer | Remote | $30-$60/hr

Codefeast • United States

On-site
USD 90,000 - 140,000