Senior ML Research Engineer - Coding-Agent Benchmarking

24-MAG

United States

Remote

USD 600,000 - 1,300,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

24-MAG LLC. seeks experienced technical professionals to advance evaluation and development of frontier coding agents.

The role sits at the intersection of AI research, software engineering, and model evaluation, designing benchmarks, datasets, and technical systems that gauge and improve coding models. The ideal candidate has a strong software-engineering background, 3+ years in ML/AI research or evaluation, and experience crafting benchmarks and evaluation methodologies.

Qualifications

  • Strong software-engineering background with Python/C++.
  • At least 3 years of experience in ML/AI research or evaluation.
  • Experience designing/validating benchmarks and assessment methodologies.

Responsibilities

  • Design and own evaluation frameworks for advanced coding agents.
  • Develop benchmark specifications, scoring rubrics, and quality standards.
  • Measure coding-model performance across diverse software tasks.
  • Analyze incorrect reasoning, tool-use failures, and execution gaps.
  • Build research tooling and collaborate with researchers and engineers.

Skills

Python
C++
Software engineering
Machine learning
Model evaluation
Benchmark design
Automation & tooling
Communication
Analytical skills
Research adaptability

Tools

Research tooling
Experimentation pipelines

Job description

24-MAG LLC. seeks experienced technical professionals to advance evaluation and development of frontier coding agents.

The role sits at the intersection of AI research, software engineering, and model evaluation, designing benchmarks, datasets, and technical systems that gauge and improve coding models. The ideal candidate has a strong software-engineering background, 3+ years in ML/AI research or evaluation, and experience crafting benchmarks and evaluation methodologies.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Coding Research Evaluation Engineer
Coding Research Evaluation Engineer

Remotebridge • United States

Remote
USD 200,000 - 260,000
Equity compensation
Bonus eligibility
Health-insurance reimbursement
+2
Remote | Member of Technical Staff, Coding Research — $600,000–$1,300,000/year
Remote | Member of Technical Staff, Coding Research — $600,000–$1,300,000/year

24-MAG • United States

Remote
USD 600,000 - 1,300,000
Remote Coding Research Engineer - Frontier AI Benchmarks
Remote Coding Research Engineer - Frontier AI Benchmarks

Pro Integrate LLC • New York (NY)

Remote
USD 200,000 - 260,000
Equity compensation
Bonuses
Health insurance premium reimbursement
+3
Senior AI Product Manager, Coding Benchmarks & Data
Senior AI Product Manager, Coding Benchmarks & Data

Scale AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff, Coding Research
Member of Technical Staff, Coding Research

Pro Integrate LLC • New York (NY)

Remote
USD 200,000 - 260,000
Equity compensation
Bonuses
Health insurance premium reimbursement
+3
Frontier ML Engineer: Benchmark AI Coding Models
Frontier ML Engineer: Benchmark AI Coding Models

Obsidian • New York (NY)

On-site
USD 207,000 - 303,000
Remote ML Benchmark Engineer: Evaluation & Experiments
Remote ML Benchmark Engineer: Evaluation & Experiments

Weekday AI • United States

Remote
USD 83,000 - 124,000
Frontier ML Engineer: AI Coding Agent Evaluator
Frontier ML Engineer: AI Coding Agent Evaluator

Mercor • New York (NY)

Hybrid
USD 275,520 - 826,560
Remote Senior Software Engineer: AI Code Evaluation
Remote Senior Software Engineer: AI Code Evaluation

24-Mag Llc • United States

Remote
USD 14,000 - 55,000
Fully remote
Flexible hours
Contract-based
Evaluation Engineer: AI Coding Benchmarks & Tests
Evaluation Engineer: AI Coding Benchmarks & Tests

Mercor • United States

Remote
USD 120,000 - 180,000