AI Evaluation & QA Analyst - Remote LATAM

Gramian Consulting Group

United States

Remote

USD 55,000 - 83,000

Part time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Gramian Consulting Group is seeking an AI Evaluation & Quality Assurance Specialist to review task quality, instructions, and scoring rubrics for AI agents. You will inspect reference solutions, grading logic, execution traces, and generated deliverables to identify defects and unfair penalties.

Comfortable reading Python, SQL, and shell scripts; passable to highly accurate analysis of technical workflows and evaluation logic is required.

Qualifications

  • Proven ability to read and interpret code, logs, and data outputs.
  • Experience evaluating AI tasks, rubrics, and reference solutions.
  • Strong analytical reasoning and ability to identify defects or biases.

Responsibilities

  • Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
  • Review AI agent execution traces, tool calls, and deliverables to determine whether outcomes are justified.
  • Audit grading logic to identify brittle checks, incorrect expected answers, unsupported rubric criteria, and unfair penalties for valid alternative solutions.
  • Investigate discrepancies between model performance, grader results, and expected outcomes.
  • Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures.
  • Independently assess automated QC findings rather than accepting them without verification.
  • Document concise, evidence-backed findings and provide actionable, reproducible feedback.
  • Flag uncertainty and verify that implemented revisions resolve previously identified issues.
  • Comfortable reading and interpreting Python, SQL, shell scripts, structured data, and execution logs.
  • Experience reviewing technical workflows, software behavior, data outputs, or evaluation logic.
  • Strong analytical skills, including the ability to verify calculations and reconcile conflicting evidence.
  • Demonstrated ability to assess the correctness and completeness of technical deliverables.
  • Strong written English with experience providing clear, specific, and reproducible feedback.
  • High attention to detail when identifying inconsistencies, missing information, and evaluation defects.

Skills

Python
SQL
Shell scripting
Data analysis
QA/testing
AI evaluation
English writing

Tools

Python
SQL
Shell

Job description

Gramian Consulting Group is seeking an AI Evaluation & Quality Assurance Specialist to review task quality, instructions, and scoring rubrics for AI agents. You will inspect reference solutions, grading logic, execution traces, and generated deliverables to identify defects and unfair penalties.

Comfortable reading Python, SQL, and shell scripts; passable to highly accurate analysis of technical workflows and evaluation logic is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Quality & Evaluation Specialist
Remote AI Quality & Evaluation Specialist

Gramian Consulting Group • United States

Remote
USD 28,000 - 55,000
Technical AI Evaluation Analyst (LATAM)
Technical AI Evaluation Analyst (LATAM)

Gramian Consulting Group • United States

Remote
USD 55,000 - 83,000
AI Quality Assurance & Evaluation Specialist
AI Quality Assurance & Evaluation Specialist

Gramian Consulting Group • United States

Remote
USD 28,000 - 55,000
AI Evaluation Specialist - Remote
AI Evaluation Specialist - Remote

YO AI Labs • New York (NY)

Remote
USD 34,000 - 55,000
AI Evaluation Specialist
AI Evaluation Specialist

micro1 • United States

On-site
AUD 70,000 - 110,000
AI Evaluation Expert - Remote
AI Evaluation Expert - Remote

YO AI Labs • New York (NY)

Remote
USD 34,000 - 76,000
Remote QA Engineer for AI Testing & Quality
Remote QA Engineer for AI Testing & Quality

YO AI Labs • Town of Texas (WI)

Remote
USD 55,000 - 83,000
Remote Code Quality Engineer for AI Evaluation
Remote Code Quality Engineer for AI Evaluation

OpenTrain AI, Inc. • United States

Remote
USD 34,000 - 69,000
AI Evaluation Expert — Remote Quality & Feedback
AI Evaluation Expert — Remote Quality & Feedback

YO AI Labs • New York (NY)

Remote
USD 34,000 - 76,000
AI QA & Competitive Intelligence Evaluator (Remote)
AI QA & Competitive Intelligence Evaluator (Remote)

Mercor • New York (NY)

Remote
USD 55,000 - 83,000