Remote AI Benchmark QA Engineer

24-Mag Llc

New York (NY)

Remote

USD 76,000 - 117,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

24-MAG LLC is seeking an experienced QA/test engineer to design robust benchmark test cases, review complex tasks, and debug Python environments for frontier AI evaluation work. The role is fully remote within the United States and requires strong attention to detail and independent problem-solving.

Your work will help ensure accurate, reproducible, and defensible evaluation results across multi-step AI tasks, with close collaboration to researchers and task authors.

Qualifications

  • At least 1 year of experience in test engineering, quality assurance, software engineering, or related technical roles.
  • Experience designing test cases and quality-review processes.
  • Strong debugging skills across complex technical systems.
  • Working proficiency in Python and Git.
  • Ability to work independently on open-ended technical problems.

Responsibilities

  • Create comprehensive test cases confirming benchmark tasks function as intended.
  • Review complex multi-step tasks and reference solutions for ambiguity and gaps.
  • Diagnose failures in Python scripts, test harnesses, and environments.
  • Develop repeatable quality processes and checklists for benchmark integrity.
  • Collaborate with researchers and task authors to improve evaluation quality.

Skills

Test design
End-to-end debugging
Quality assurance
Attention to detail

Education

Master's degree in STEM
Equivalent practical experience

Tools

Python
Git

Job description

24-MAG LLC is seeking an experienced QA/test engineer to design robust benchmark test cases, review complex tasks, and debug Python environments for frontier AI evaluation work. The role is fully remote within the United States and requires strong attention to detail and independent problem-solving.

Your work will help ensure accurate, reproducible, and defensible evaluation results across multi-step AI tasks, with close collaboration to researchers and task authors.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote QA/Test Engineer for AI Benchmarks
Remote QA/Test Engineer for AI Benchmarks

Dorado • United States

Remote
USD 90,000 - 130,000
Remote AI QA Engineer — Benchmark & Test Specialist
Remote AI QA Engineer — Benchmark & Test Specialist

Mercor • United States

On-site
Remote Python Engineer for AI Benchmarking
Remote Python Engineer for AI Benchmarking

24-Mag Llc • New York (NY)

Remote
Remote QA Engineer for AI Benchmarking & Quality Review
Remote QA Engineer for AI Benchmarking & Quality Review

Mercor • United States

On-site
Remote CS Research Expert - Benchmark & ML Systems
Remote CS Research Expert - Benchmark & ML Systems

24-Mag Llc • New York (NY)

Remote
USD 76,000 - 103,000
Remote work
Flexible schedule
Competitive hourly rate
Remote Part-Time Senior Engineering Assessment Consultant
Remote Part-Time Senior Engineering Assessment Consultant

24-Mag Llc • New York (NY)

Remote
USD 69,000 - 96,000
Remote AI QA Engineer - AI-Powered Automated Testing
Remote AI QA Engineer - AI-Powered Automated Testing

Public Partnerships LLC • United States

On-site
USD 110,000 - 140,000
Remote QA Engineer – AI Platform & Test Automation
Remote QA Engineer – AI Platform & Test Automation

WebSenor Ltd • United States

Remote
USD 70,000 - 110,000
Remote Data Scientist & Quant Analyst for AI Evaluation
Remote Data Scientist & Quant Analyst for AI Evaluation

24-Mag Llc • New York (NY)

Remote
Remote Physics AI Benchmark Designer
Remote Physics AI Benchmark Designer

United States Digital Space LLC • Germany (OH)

On-site
USD 45,000 - 65,000