Frontier AI Benchmark Researcher — Remote

Synthires

United States

Remote

USD 193,000 - 207,000

Full time

8 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Fully remote opportunity
Flexible workload (5–40 hours/week)
Fully asynchronous work environment

Job summary

AfterQuery is seeking experienced Machine Learning Research Experts with a strong research track record to design benchmark tasks, evaluate model capabilities, and contribute expert-level ML reasoning to next-generation AI systems. The role emphasizes designing rigorous evaluation tasks and may focus on areas such as Language Models, Deep Learning, RL, CV, Generative AI, and more.

Fully remote, rolling applications, and asynchronous work are offered.

Qualifications

  • At least one first-author research publication in ML/AI or a closely related field.
  • Master's or PhD in Machine Learning, Computer Science, Statistics, Mathematics, or related quantitative discipline (completed or in progress).
  • Minimum 1 year of hands-on machine learning research experience.
  • Strong experience developing, training, evaluating, or deploying ML models.
  • Proficiency in ML programming and experimentation.
  • Ability to design rigorous research problems and evaluation frameworks.
  • Excellent written communication and technical documentation skills.

Responsibilities

  • Design realistic ML research tasks and benchmark problems.
  • Develop expert-level reference solutions and evaluation methodologies.
  • Create grading rubrics for model reasoning, implementation quality, and scientific correctness.
  • Evaluate AI-generated solutions for technical accuracy and research validity.
  • Author problem sets covering advanced ML concepts and research workflows.
  • Review model outputs and identify weaknesses in reasoning, implementation, experimentation, and interpretation; document solutions and benchmarks.
  • Contribute domain expertise to improve frontier AI evaluation systems.

Skills

First-author publication
Hands-on ML research
ML programming
Experimentation
Written communication
Research problem design

Education

Master's or PhD in ML/CS/Statistics/Mathematics

Job description

AfterQuery is seeking experienced Machine Learning Research Experts with a strong research track record to design benchmark tasks, evaluate model capabilities, and contribute expert-level ML reasoning to next-generation AI systems. The role emphasizes designing rigorous evaluation tasks and may focus on areas such as Language Models, Deep Learning, RL, CV, Generative AI, and more.

Fully remote, rolling applications, and asynchronous work are offered.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Scientist (Remote | $140–$150/hr)
AI Research Scientist (Remote | $140–$150/hr)

Synthires • United States

Remote
USD 193,000 - 207,000
Fully remote opportunity
Flexible workload (5–40 hours/week)
Fully asynchronous work environment
Remote ML Benchmark Engineer: Evaluation & Experiments
Remote ML Benchmark Engineer: Evaluation & Experiments

Weekday AI • United States

Remote
USD 83,000 - 124,000
Remote STEM Researcher - AI Evaluation & Benchmark Design
Remote STEM Researcher - AI Evaluation & Benchmark Design

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Applied AI Research Engineer (Fully Remote)
Applied AI Research Engineer (Fully Remote)

Appen Limited • United States

Remote
USD 120,000 - 180,000
Staff Research Engineer, Frontier AI Evaluation Benchmarks
Staff Research Engineer, Frontier AI Evaluation Benchmarks

OpenTrain AI, Inc. • United States

Remote
USD 103,000 - 193,000
Frontier AI Infrastructure Research Scientist
Frontier AI Infrastructure Research Scientist

Greylock Partners • New York (NY)

On-site
USD 180,000 - 240,000
Staff Research Engineer, Frontier AI & Evaluation (Remote)
Staff Research Engineer, Frontier AI & Evaluation (Remote)

Crossing Hurdles • United States

Remote
USD 48,000 - 103,000
Staff Research Engineer ( Frontier AI, RL & Evaluation)
Staff Research Engineer ( Frontier AI, RL & Evaluation)

Crossing Hurdles • United States

Remote
USD 150,000 - 210,000
Machine Learning Engineer - Model Evaluation & Experimentation
Machine Learning Engineer - Model Evaluation & Experimentation

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Remote ML Benchmark Engineer & Evaluation Specialist
Remote ML Benchmark Engineer & Evaluation Specialist

Weekday 1 • United States

Remote
USD 83,000 - 124,000