Technical AI Evaluation Analyst (LATAM)

Gramian Consulting Group

United States

Remote

USD 55,000 - 83,000

Part time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Gramian Consulting Group is seeking an AI Evaluation & Quality Assurance Specialist to review task quality, instructions, and scoring rubrics for AI agents. You will inspect reference solutions, grading logic, execution traces, and generated deliverables to identify defects and unfair penalties.

Comfortable reading Python, SQL, and shell scripts; passable to highly accurate analysis of technical workflows and evaluation logic is required.

Qualifications

  • Proven ability to read and interpret code, logs, and data outputs.
  • Experience evaluating AI tasks, rubrics, and reference solutions.
  • Strong analytical reasoning and ability to identify defects or biases.

Responsibilities

  • Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
  • Review AI agent execution traces, tool calls, and deliverables to determine whether outcomes are justified.
  • Audit grading logic to identify brittle checks, incorrect expected answers, unsupported rubric criteria, and unfair penalties for valid alternative solutions.
  • Investigate discrepancies between model performance, grader results, and expected outcomes.
  • Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures.
  • Independently assess automated QC findings rather than accepting them without verification.
  • Document concise, evidence-backed findings and provide actionable, reproducible feedback.
  • Flag uncertainty and verify that implemented revisions resolve previously identified issues.
  • Comfortable reading and interpreting Python, SQL, shell scripts, structured data, and execution logs.
  • Experience reviewing technical workflows, software behavior, data outputs, or evaluation logic.
  • Strong analytical skills, including the ability to verify calculations and reconcile conflicting evidence.
  • Demonstrated ability to assess the correctness and completeness of technical deliverables.
  • Strong written English with experience providing clear, specific, and reproducible feedback.
  • High attention to detail when identifying inconsistencies, missing information, and evaluation defects.

Skills

Python
SQL
Shell scripting
Data analysis
QA/testing
AI evaluation
English writing

Tools

Python
SQL
Shell

Job description

About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.


About the Role

We are seeking a detail-oriented AI Evaluation & Quality Assurance Specialist to review the quality, correctness, and fairness of tasks designed to evaluate AI agents. You will inspect reference solutions, grading logic, execution traces, and generated deliverables to identify task defects, evaluation errors, and unjustified model failures. This role requires strong technical fluency, independent analytical judgment, and the ability to produce clear, evidence-based feedback.


LOCATION: Remote – Latin America (LATAM)


CONTRACT: Hourly Contractor


COMMITMENT: 40 hours per week


TIME OVERLAP: 8 hours of mandatory PST overlap


DURATION: 10 weeks


START DATE: Immediately


Key Responsibilities


  • Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.

  • Review AI agent execution traces, tool calls, and deliverables to determine whether outcomes are justified.

  • Audit grading logic to identify brittle checks, incorrect expected answers, unsupported rubric criteria, and unfair penalties for valid alternative solutions.

  • Investigate discrepancies between model performance, grader results, and expected outcomes.

  • Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures.

  • Independently assess automated QC findings rather than accepting them without verification.

  • Document concise, evidence-backed findings and provide actionable, reproducible feedback.

  • Flag uncertainty and verify that implemented revisions resolve previously identified issues.



  • Comfortable reading and interpreting Python, SQL, shell scripts, structured data, and execution logs.



  • Experience reviewing technical workflows, software behavior, data outputs, or evaluation logic.

  • Strong analytical skills, including the ability to verify calculations and reconcile conflicting evidence.

  • Demonstrated ability to assess the correctness and completeness of technical deliverables.

  • Strong written English with experience providing clear, specific, and reproducible feedback.

  • High attention to detail when identifying inconsistencies, missing information, and evaluation defects.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Quality Assurance & Evaluation Specialist
AI Quality Assurance & Evaluation Specialist

Gramian Consulting Group • United States

Remote
USD 28,000 - 55,000
Remote AI Quality & Evaluation Specialist
Remote AI Quality & Evaluation Specialist

Gramian Consulting Group • United States

Remote
USD 28,000 - 55,000
AI Evaluation & QA Analyst - Remote LATAM
AI Evaluation & QA Analyst - Remote LATAM

Gramian Consulting Group • United States

Remote
USD 55,000 - 83,000
AI Evaluation Specialist
AI Evaluation Specialist

micro1 • United States

On-site
AUD 70,000 - 110,000
Technical Program Operations Lead - AI Engineering
Technical Program Operations Lead - AI Engineering

Jobs in JS • United States

On-site
USD 150,000 - 190,000
Software Engineer - Backend & AI Code Evaluation
Software Engineer - Backend & AI Code Evaluation

Gramian Consulting Group • United States

Remote
USD 120,000 - 170,000
Life Sciences Researcher (AI Evaluation)
Life Sciences Researcher (AI Evaluation)

Gramian Consulting • United States

Remote
PKR 1,378,000 - 2,755,000
AI Evaluation Specialist - Remote
AI Evaluation Specialist - Remote

YO AI Labs • New York (NY)

Remote
USD 34,000 - 55,000
AI Evaluation Specialist
AI Evaluation Specialist

Weekday 1 • United States

Remote
USD 80,000 - 113,000
Generative AI Evaluator | $30/hr Remote
Generative AI Evaluator | $30/hr Remote

Crossing Hurdles • United States

On-site
CAD 27,552 - 41,328