Remote QA/Test Engineer — AI Benchmark Validation

Weekday AI (YC W21)

United States

On-site

USD 83,000 - 124,000

Full time

12 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Weekday AI (YC W21) in the United States seeks an experienced QA/Test Engineer to establish quality standards and testing processes for agentic evaluation benchmarks in advanced AI models. You will design tests, review task quality, debug environments with Python, and collaborate with researchers.

This fully remote role requires approximately 35 hours per week and is offered as Full-Time employment. You will help ensure reliable, unambiguous benchmarking and rigorous evaluation of AI tasks.

Qualifications

  • MSc or PhD in STEM or equivalent research/engineering experience.
  • 1+ year in QA or test engineering with ownership of quality.
  • Strong ability to design test cases, QA processes, and end-to-end investigation.

Responsibilities

  • Design test cases for benchmark tasks, including edge cases and failures.
  • Review task quality, requirements, and reference solutions for ambiguities.
  • Debug and troubleshoot environments using Python and development tools.
  • Develop repeatable QA processes and checklists.
  • Identify evaluation gaps and ensure reliable scoring.
  • Collaborate with researchers and task authors to validate fixes.
  • Maintain high-quality standards across evaluation datasets.

Skills

Python
Git
Test design
Debugging
Documentation
Independent work
Attention to detail

Education

MSc or PhD in STEM

Job description

Weekday AI (YC W21) in the United States seeks an experienced QA/Test Engineer to establish quality standards and testing processes for agentic evaluation benchmarks in advanced AI models. You will design tests, review task quality, debug environments with Python, and collaborate with researchers.

This fully remote role requires approximately 35 hours per week and is offered as Full-Time employment. You will help ensure reliable, unambiguous benchmarking and rigorous evaluation of AI tasks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote QA Engineer for AI Validation & Automation
Remote QA Engineer for AI Validation & Automation

YO AI Labs • San Francisco (CA)

Remote
USD 83,000 - 138,000
Remote QA Engineer - AI Testing & Automation
Remote QA Engineer - AI Testing & Automation

YO AI Labs • Boston (MA)

Remote
USD 70,000 - 110,000
Remote QA Engineer for AI Testing & Quality Assurance
Remote QA Engineer for AI Testing & Quality Assurance

YO AI Labs • Houston (TX)

Remote
USD 90,000 - 130,000
Remote QA Engineer for AI Testing & Automation
Remote QA Engineer for AI Testing & Automation

YO AI Labs • San Francisco (CA)

Remote
USD 55,000 - 110,000
Remote QA Engineer for AI Testing & Quality Validation
Remote QA Engineer for AI Testing & Quality Validation

YO AI Labs • Miami (FL)

Remote
USD 55,000 - 83,000
QA/Test Engineer
QA/Test Engineer

Weekday AI (YC W21) • United States

On-site
USD 83,000 - 124,000
Remote QA Engineer for AI Testing & Quality
Remote QA Engineer for AI Testing & Quality

YO AI Labs • Miami (FL)

Remote
USD 55,000 - 110,000
Remote QA Engineer for AI & Software Testing
Remote QA Engineer for AI & Software Testing

YO AI Labs • Illinois

Remote
USD 55,000 - 83,000
Remote QA Engineer for AI Testing & Quality
Remote QA Engineer for AI Testing & Quality

YO AI Labs • Town of Texas (WI)

Remote
USD 55,000 - 83,000
Remote AI Benchmark Test Engineer
Remote AI Benchmark Test Engineer

Mercor • New York (NY)

Remote
USD 85,000 - 120,000