Remote AI Benchmark Test Engineer

Mercor

New York (NY)

A distancia

USD 85.000 - 120.000

Jornada completa

14 días+
Generador de candidaturas

Una candidatura completa en un minuto — currículum y carta de presentación adaptados, listos para enviar.

Supera los filtros ATS

Descripción de la vacante

Mercor, a leading AI platform, seeks a QA/test engineer to join its GenAI benchmarking effort. You will design checks, review tasks, and debug environments, ensuring edge cases are covered and results are robust.

This is a full-time W-2 role placed with Cincinnatus LLC, fully remote in the United States. You will work about 35 hours per week, collaborating with researchers and task authors to build repeatable quality processes and clear documentation, with independence to tackle ambiguous

Formación

  • MSc or PhD in STEM or equivalent practical experience.
  • 1+ years of experience in test engineering, quality assurance, or a research/software engineering role with strong quality ownership.
  • Demonstrated skill designing test cases and quality-review processes, and debugging complex systems end-to-end.
  • Working proficiency in Python and Git, and comfort navigating unfamiliar codebases and environments.
  • Exceptional attention to detail and clear written documentation habits.
  • Past experience in AI training, model evaluation, or quality review of AI-generated work is preferred.
  • A perfectionist mindset: creativity in finding what others missed, and the ability to work independently through ambiguous, open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.

Responsabilidades

  • Design checks: Create test cases that confirm each task works as intended — including the tricky edge cases.
  • Review tasks: Give tasks and reference solutions a careful read before they're finalized, catching ambiguity and gaps early.
  • Debug: Roll up your sleeves in Python when a task or its checks don't behave the way they should.
  • Shape the process: Help build simple, repeatable quality checklists, and share feedback authors can act on right away.
  • Protect the results: Watch for shortcuts and grading gaps in AI agent runs so benchmark scores stay trustworthy.

Conocimientos

Test engineering
Quality assurance
Design test cases
Python
Git
Documentation
Independent work
Attention to detail
AI model evaluation

Educación

MSc or PhD in STEM

Descripción del empleo

Mercor, a leading AI platform, seeks a QA/test engineer to join its GenAI benchmarking effort. You will design checks, review tasks, and debug environments, ensuring edge cases are covered and results are robust.

This is a full-time W-2 role placed with Cincinnatus LLC, fully remote in the United States. You will work about 35 hours per week, collaborating with researchers and task authors to build repeatable quality processes and clear documentation, with independence to tackle ambiguous

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Quantitative Analyst for GenAI Benchmarking
Remote Quantitative Analyst for GenAI Benchmarking

Mercor • New York (NY)

A distancia
USD 100.000 - 180.000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

Presencial
USD 120.000 - 170.000
Senior Software Domain Expert — GenAI QA & Benchmarks
Senior Software Domain Expert — GenAI QA & Benchmarks

Mercor • San Francisco (CA)

Híbrido
USD 180.000 - 240.000
Remote QA/Test Engineer — AI Benchmark Validation
Remote QA/Test Engineer — AI Benchmark Validation

Weekday AI (YC W21) • EE. UU.

Presencial
USD 83.000 - 124.000
GenAI Benchmark Research Scientist - Remote (35h/wk)
GenAI Benchmark Research Scientist - Remote (35h/wk)

Mercor • San Francisco (CA)

A distancia
USD 120.000 - 180.000
QA Engineer - AI Benchmarking (Remote, 35h/wk)
QA Engineer - AI Benchmarking (Remote, 35h/wk)

Weekday AI • EE. UU.

A distancia
USD 83.000 - 124.000
AI Benchmark Engineer — PhD (Remote & Flexible)
AI Benchmark Engineer — PhD (Remote & Flexible)

Mercor • New York (NY)

Presencial
USD 55.000 - 110.000
Remote Materials Science AI Benchmark Engineer
Remote Materials Science AI Benchmark Engineer

Weekday AI • EE. UU.

A distancia
USD 83.000 - 110.000
Remote AI Benchmarking Consultant (Part-Time)
Remote AI Benchmarking Consultant (Part-Time)

Obsidian • San Francisco (CA)

Presencial
USD 27.000 - 55.000
Senior AI Domain Expert — GenAI Benchmarks & Specs
Senior AI Domain Expert — GenAI Benchmarks & Specs

Mercor • San Francisco (CA)

Híbrido
USD 180.000 - 250.000