QA/Test Engineer

Dorado

United States

Remote

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cincinnatus LLC is hiring a QA/Test Engineer to help design and review complex AI evaluation tasks. You will run tasks, probe edge cases, and debug environments in a remote US role around 35 hours per week. Collaboration with researchers and task authors is a core part of the job.

The ideal candidate has MSc/PhD or equivalent research/engineering experience, strong Python and Git skills, and a keen eye for detail. Prior AI training or model evaluation experience is a plus.

Qualifications

  • MSc or PhD in a STEM field, or equivalent practical experience.
  • 1+ years of experience in test engineering, QA, or a research/software engineering role with strong quality ownership.
  • Designing test cases and quality-review processes, and debugging complex systems end-to-end.
  • Working proficiency in Python and Git, and comfort navigating unfamiliar codebases.
  • Exceptional attention to detail and clear written documentation habits.
  • Past experience in AI training, model evaluation, or quality review of AI-generated work is preferred.

Responsibilities

  • Design checks: create test cases with edge cases.
  • Review tasks and reference solutions for ambiguity.
  • Debug tasks using Python.
  • Shape repeatable quality checklists and give actionable feedback.
  • Protect results by identifying shortcuts and grading gaps.

Skills

Test engineering
Python
Git
Quality ownership
Attention to detail
Independent work

Education

MSc/PhD in STEM

Job description

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.

1. Overview

A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models, and complex multi-step tasks are only useful if they are airtight: unambiguous, correctly graded, and robust to shortcuts. We are seeking experienced QA and test engineers to be the quality backbone of this benchmark — designing the test cases and review processes that guarantee every task measures what it claims to measure.

Each task under review represents one to two days of expert effort and spans multiple technical skills, so quality review here means genuinely understanding the task: running it, probing its edge cases, and debugging its environment. You will work in a tight feedback loop with the lab's researchers and task authors.

This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week.

2. Key Responsibilities
  • Design checks: Create test cases that confirm each task works as intended — including the tricky edge cases.

  • Review tasks: Give tasks and reference solutions a careful read before they're finalized, catching ambiguity and gaps early.

  • Debug: Roll up your sleeves in Python when a task or its checks don't behave the way they should.

  • Shape the process: Help build simple, repeatable quality checklists, and share feedback authors can act on right away.

  • Protect the results: Watch for shortcuts and grading gaps in AI agent runs so benchmark scores stay trustworthy.

3. Core Qualifications
  • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain.

  • 1+ years of experience in test engineering, quality assurance, or a research/software engineering role with strong quality ownership.

  • Demonstrated skill designing test cases and quality-review processes, and debugging complex systems end-to-end.

  • Working proficiency in Python and Git, and comfort navigating unfamiliar codebases and environments.

  • Exceptional attention to detail and clear written documentation habits.

  • Past experience in AI training, model evaluation, or quality review of AI-generated work is preferred.

  • A perfectionist mindset: creativity in finding what others missed, and the ability to work independently through ambiguous, open-ended problems.

  • Ability to engage reliably for approximately 35 hours per week.

About Cincinnatus LLC

Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.

Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows.

Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC.

Equal Employment Opportunity

Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Dorado • United States

Remote
USD 90,000 - 120,000
Data Science Expert - AI/ML
Data Science Expert - AI/ML

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
Engineering & Software Domain Expert Mercor · Bay Area, CA $65-105/hr →
Engineering & Software Domain Expert Mercor · Bay Area, CA $65-105/hr →

Dorado • California (MO), Northern (KY)

Hybrid
USD 180,000 - 240,000
Software Domain Expert
Software Domain Expert

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Machine Learning Engineer — Model Evaluation & Experimentation
Machine Learning Engineer — Model Evaluation & Experimentation

Dorado • United States

Remote
USD 120,000 - 180,000
STEM Researcher — Computational Fields
STEM Researcher — Computational Fields

Dorado • United States

Remote
USD 105,000 - 150,000
Quality Engineer (AI & Test Automation)
Quality Engineer (AI & Test Automation)

Cognizant • Hartford (CT)

On-site
USD 58,500 - 71,500
Medical/Dental/Vision/Life Insurance
Paid holidays plus Paid Time Off
401(k) plan and contributions
+3
Quality Engineer (AI & Test Automation)
Quality Engineer (AI & Test Automation)

Cognizant • Richmond (VA)

Hybrid
USD 58,500 - 71,500
Medical/Dental/Vision Insurance
Paid Time Off
401(k) plan
+1
Quality Engineer (AI & Test Automation)
Quality Engineer (AI & Test Automation)

Cognizant • Columbus (OH)

Hybrid
USD 58,500 - 71,500
Medical/Dental/Vision/Life Insurance
Paid holidays and Paid Time Off
401(k) plan and contributions
+3
LLM Red Team Specialist — Failure Modes & Edge Cases
LLM Red Team Specialist — Failure Modes & Edge Cases

Dorado • United States

Remote
USD 120,000 - 180,000