QA Engineer - AI Benchmarking (Remote, 35h/wk)

Weekday AI

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Weekday AI is seeking an experienced QA and Test Engineer for a fully remote, independent contractor role in the United States. The position focuses on building reliable, reproducible AI evaluation benchmarks and validating complex evaluation tasks with strong debugging of Python-based scripts.

You will design test cases, review tasks for ambiguity, and collaborate with AI researchers to ensure benchmark integrity and rigorous testing methodologies. approx.

Qualifications

  • Master's degree, PhD, or equivalent practical experience in a STEM discipline involving software engineering, research, or advanced technical problem solving.
  • Minimum 1 year of professional experience in Quality Assurance, Test Engineering, Software Engineering, Research Engineering, or related field with strong quality ownership.
  • Proven experience designing test cases, validating complex software systems, and debugging end-to-end workflows.
  • Strong proficiency in Python and Git, with the ability to troubleshoot unfamiliar codebases and technical environments.
  • Excellent analytical thinking, problem-solving skills, and exceptional attention to detail.
  • Experience documenting bugs, test strategies, and technical findings with clear written communication.
  • Experience evaluating AI systems, machine learning models, or AI-generated outputs is preferred.
  • Ability to work independently while managing multiple complex tasks with minimal supervision.
  • Ability to commit approximately 35 hours per week on a consistent basis.

Responsibilities

  • Design comprehensive test cases that validate evaluation tasks, including complex edge cases and unexpected scenarios.
  • Review benchmark tasks and reference solutions to identify ambiguity, inconsistencies, missing requirements, and grading gaps.
  • Debug task environments and Python-based evaluation scripts to ensure reliable execution and accurate results.
  • Develop repeatable quality assurance processes, validation checklists, and testing frameworks for benchmark creation.
  • Identify potential shortcuts, exploits, or weaknesses that could compromise evaluation accuracy or benchmark integrity.
  • Collaborate with AI researchers, engineers, and task authors to improve task quality, reproducibility, and technical rigor.

Skills

Python
Git
Test design
Edge case analysis

Education

Master's degree or PhD in STEM

Job description

Weekday AI is seeking an experienced QA and Test Engineer for a fully remote, independent contractor role in the United States. The position focuses on building reliable, reproducible AI evaluation benchmarks and validating complex evaluation tasks with strong debugging of Python-based scripts.

You will design test cases, review tasks for ambiguity, and collaborate with AI researchers to ensure benchmark integrity and rigorous testing methodologies. approx.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote QA/Test Engineer — AI Benchmark Validation
Remote QA/Test Engineer — AI Benchmark Validation

Weekday AI (YC W21) • United States

On-site
USD 83,000 - 124,000
AI Benchmark Engineer (Python) — Remote
AI Benchmark Engineer (Python) — Remote

Weekday AI • United States

Remote
USD 83,000 - 124,000
QA/Test Engineer
QA/Test Engineer

Weekday AI • United States

Remote
USD 83,000 - 124,000
AI Benchmark Architect (Remote) | Python
AI Benchmark Architect (Remote) | Python

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
QA/Test Engineer
QA/Test Engineer

Weekday AI (YC W21) • United States

On-site
USD 83,000 - 124,000
AI Benchmark Designer (Remote)
AI Benchmark Designer (Remote)

Weekday AI • United States

Remote
USD 83,000 - 103,000
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
Remote AI Benchmark Test Engineer
Remote AI Benchmark Test Engineer

Mercor • New York (NY)

Remote
USD 85,000 - 120,000
Remote AI Benchmark Review & QA Specialist
Remote AI Benchmark Review & QA Specialist

AI Trainer Jobs • United States

Remote
USD 91,000 - 116,000
Remote Applied Engineering Benchmark QA Specialist
Remote Applied Engineering Benchmark QA Specialist

AI Trainer Jobs • United States

Remote
USD 84,000 - 106,000