Remote AI Benchmark Test Engineer

Mercor

New York (NY)

Remote

USD 85,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor, a leading AI platform, seeks a QA/test engineer to join its GenAI benchmarking effort. You will design checks, review tasks, and debug environments, ensuring edge cases are covered and results are robust.

This is a full-time W-2 role placed with Cincinnatus LLC, fully remote in the United States. You will work about 35 hours per week, collaborating with researchers and task authors to build repeatable quality processes and clear documentation, with independence to tackle ambiguous

Qualifications

  • MSc or PhD in STEM or equivalent practical experience.
  • 1+ years of experience in test engineering, quality assurance, or a research/software engineering role with strong quality ownership.
  • Demonstrated skill designing test cases and quality-review processes, and debugging complex systems end-to-end.
  • Working proficiency in Python and Git, and comfort navigating unfamiliar codebases and environments.
  • Exceptional attention to detail and clear written documentation habits.
  • Past experience in AI training, model evaluation, or quality review of AI-generated work is preferred.
  • A perfectionist mindset: creativity in finding what others missed, and the ability to work independently through ambiguous, open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.

Responsibilities

  • Design checks: Create test cases that confirm each task works as intended — including the tricky edge cases.
  • Review tasks: Give tasks and reference solutions a careful read before they're finalized, catching ambiguity and gaps early.
  • Debug: Roll up your sleeves in Python when a task or its checks don't behave the way they should.
  • Shape the process: Help build simple, repeatable quality checklists, and share feedback authors can act on right away.
  • Protect the results: Watch for shortcuts and grading gaps in AI agent runs so benchmark scores stay trustworthy.

Skills

Test engineering
Quality assurance
Design test cases
Python
Git
Documentation
Independent work
Attention to detail
AI model evaluation

Education

MSc or PhD in STEM

Job description

Mercor, a leading AI platform, seeks a QA/test engineer to join its GenAI benchmarking effort. You will design checks, review tasks, and debug environments, ensuring edge cases are covered and results are robust.

This is a full-time W-2 role placed with Cincinnatus LLC, fully remote in the United States. You will work about 35 hours per week, collaborating with researchers and task authors to build repeatable quality processes and clear documentation, with independence to tackle ambiguous

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote QA Engineer for AI Benchmarking & Quality Review
Remote QA Engineer for AI Benchmarking & Quality Review

Mercor • United States

On-site
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
Remote Quantitative Analyst for GenAI Benchmarking
Remote Quantitative Analyst for GenAI Benchmarking

Mercor • New York (NY)

Remote
USD 100,000 - 180,000
Remote QA/Test Engineer for AI Benchmarks
Remote QA/Test Engineer for AI Benchmarks

Dorado • United States

Remote
USD 90,000 - 130,000
Senior Software Domain Expert — GenAI QA & Benchmarks
Senior Software Domain Expert — GenAI QA & Benchmarks

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Remote AI Benchmark QA Engineer
Remote AI Benchmark QA Engineer

24-Mag Llc • New York (NY)

Remote
USD 76,000 - 117,000
QA/Test Engineer
QA/Test Engineer

Dorado • United States

Remote
USD 90,000 - 130,000
GenAI Benchmark Research Scientist - Remote (35h/wk)
GenAI Benchmark Research Scientist - Remote (35h/wk)

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
AI Benchmark Engineer — PhD (Remote & Flexible)
AI Benchmark Engineer — PhD (Remote & Flexible)

Mercor • New York (NY)

On-site
USD 55,000 - 110,000
Remote AI Benchmarking Consultant (Part-Time)
Remote AI Benchmarking Consultant (Part-Time)

Mercor • San Francisco (CA)

On-site