AI QA Trainer - LLM Evaluation - Freelance Project

Meridial

United States

Remote

USD 8,265 - 89,544

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Secure computer and high-speed internet required

Job summary

Meridial seeks an AI QA Trainer for a remote freelance project to evaluate and enhance the performance of large-scale language models. The ideal candidate will possess expertise in model evaluation and exposure to QA for ML/AI systems. Responsibilities include verifying model outputs, designing test plans, and documenting error modes.

This position offers a pay range of $6–$65 per hour, with the rate determined by experience and expertise.

Qualifications

  • Experience with ML/AI systems QA.
  • Expertise in adversarial testing/red-teaming.
  • Experience with test automation frameworks.

Responsibilities

  • Converse with models on evaluation prompts.
  • Verify factual accuracy and logical soundness.
  • Document and report failures effectively.

Skills

Model evaluation
LLM safety
Data quality assurance
Adversarial testing
Regression testing

Education

Bachelor’s, Master’s, or PhD in relevant fields

Tools

Python
SQL
PyTest

Job description

AI QA Trainer - LLM Evaluation - Freelance Project

World Wide - Remote

Are you an AI QA expert eager to shape the future of AI? Large-scale language models are evolving from clever chatbots into enterprise-grade platforms. With rigorous evaluation data, tomorrow’s AI can democratize world-class education, keep pace with cutting-edge research, and streamline workflows for teams everywhere. That quality begins with you—we need your expertise to harden model reasoning and reliability.

We’re looking for AI QA trainers who live and breathe model evaluation, LLM safety, prompt robustness, data quality assurance, multilingual and domain-specific testing, grounding verification, and compliance/readiness checks. You’ll challenge advanced language models on tasks like hallucination detection, factual consistency, prompt-injection and jailbreak resistance, bias/fairness audits, chain-of-reasoning reliability, tool-use correctness, retrieval-augmentation fidelity, and end-to-end workflow validation—documenting every failure mode so we can raise the bar.

On a typical day, you will converse with the model on real-world scenarios and evaluation prompts, verify factual accuracy and logical soundness, design and run test plans and regression suites, build clear rubrics and pass/fail criteria, capture reproducible error traces with root-cause hypotheses, and suggest improvements to prompt engineering, guardrails, and evaluation metrics (e.g., precision/recall, faithfulness, toxicity, and latency SLOs). You’ll also partner on adversarial red-teaming, automation (Python/SQL), and dashboarding to track quality deltas over time.

A bachelor’s, master’s, or PhD in computer science, data science, computational linguistics, statistics, or a related field is ideal; shipped QA for ML/AI systems, safety/red-team experience, test automation frameworks (e.g., PyTest), and hands-on work with LLM eval tooling (e.g., OpenAI Evals, RAG evaluators, W&B) signal fit. Skills that stand out include evaluation rubric design, adversarial testing/red-teaming, regression testing at scale, bias/fairness auditing, grounding verification, prompt and system-prompt engineering, test automation (Python/SQL), and high-signal bug reporting. Clear, metacognitive communication—“showing your work”—is essential.

We offer a pay range of $6–$65 per hour, with the exact rate determined after evaluating your experience, expertise, and geographic location. As a contractor you’ll supply a secure computer and high-speed internet; company-sponsored benefits such as health insurance and PTO do not apply.

Employment type: Contract | Workplace type: Remote | Seniority level: Mid-Senior Level

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI QA Trainer & Evaluation Specialist
Remote AI QA Trainer & Evaluation Specialist

Meridial • United States

Remote
Secure computer and high-speed internet required
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Austin (TX)

Remote
Competitive pay
Flexible schedule
Experience in advanced AI projects
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Generative AI Evaluator | $30/hr Remote
Generative AI Evaluator | $30/hr Remote

Crossing Hurdles • United States

Remote
Evaluation Scenario Writer - QA
Evaluation Scenario Writer - QA

Mindrift • New York (NY)

Remote
Competitive pay up to $55/hour
Flexible remote work
Experience with advanced AI projects
AI Agent Evaluation Analyst
AI Agent Evaluation Analyst

Mindrift • Dallas (TX)

Remote
Flexible remote work
Competitive pay up to $55/hour
Experience in advanced AI projects
AI QA Engineer
AI QA Engineer

Cavendish Professionals • Town of Italy (NY)

On-site
USD 95,000 - 120,000
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Alabama

Remote
Flexible working hours
Competitive pay up to $80/hour
Experience in advanced AI projects
Training and Development Specialist - Freelance AI Trainer Project
Training and Development Specialist - Freelance AI Trainer Project

Meridial • United States

On-site
Flexible working hours
Opportunity to shape AI training
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Mississippi

Remote
USD 55,000 - 75,000
Competitive pay up to $80/hr
Flexible, remote work
Participation in advanced AI projects