Remote Coding Research Engineer - Frontier AI Benchmarks

Pro Integrate LLC

New York (NY)

Remote

USD 200,000 - 260,000

Full time

13 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity compensation
Bonuses
Health insurance premium reimbursement
Paid time off
401(k) with company match
Remote-first work environment

Job summary

Pro Integrate LLC is seeking a Member of Technical Staff to advance the evaluation and development of frontier coding agents. You will work at the intersection of AI research, software engineering, and model evaluation to design benchmarks, data systems, and experiments that shape how coding models are measured and improved.

You will collaborate with researchers, engineers, and applied AI teams to design experiments, measure model capabilities, and communicate findings through technical reports

Qualifications

  • Strong software engineering background with Python or C++.
  • 3+ years of experience in software engineering, ML, AI research, evaluation, or related disciplines.
  • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
  • Familiarity with LLMs, coding agents, RL, model evaluation, or related AI systems.
  • Proven ability to build tooling, automate workflows, and improve technical processes.
  • Strong analytical skills to investigate model behavior and derive insights.
  • Excellent written and verbal communication to convey findings.
  • Comfortable in fast-moving research environments with ambiguity.

Responsibilities

  • Design and own evaluation frameworks for coding agents, including benchmarks and scoring rubrics.
  • Lead end-to-end research initiatives measuring coding model performance.
  • Develop high-quality datasets, golden examples, and evaluation protocols.
  • Analyze model behavior and identify weaknesses for improvements in training/evaluation.
  • Build tooling and infrastructure for large-scale experimentation and review workflows.
  • Establish best practices for coding-agent assessment, ensuring rigor and reproducibility.
  • Partner with researchers, engineers, and applied AI teams to design experiments.
  • Contribute to technical reports and benchmark studies.

Skills

Python
C++
Software engineering
ML/AI research
Benchmark design
Experimentation
Analytical skills
Communication

Tools

Git
CI/CD

Job description

Pro Integrate LLC is seeking a Member of Technical Staff to advance the evaluation and development of frontier coding agents. You will work at the intersection of AI research, software engineering, and model evaluation to design benchmarks, data systems, and experiments that shape how coding models are measured and improved.

You will collaborate with researchers, engineers, and applied AI teams to design experiments, measure model capabilities, and communicate findings through technical reports

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Coding Research Evaluation Engineer
Coding Research Evaluation Engineer

Remotebridge • United States

Remote
USD 200,000 - 260,000
Equity compensation
Bonus eligibility
Health-insurance reimbursement
+2
AI Coding Data Engineer for Frontier Model Evaluation
AI Coding Data Engineer for Frontier Model Evaluation

Mercor • New York (NY)

On-site
USD 454,608 - 647,472
Frontier ML Engineer: Benchmark AI Coding Models
Frontier ML Engineer: Benchmark AI Coding Models

Obsidian • New York (NY)

On-site
USD 207,000 - 303,000
Backend Engineer: AI Coding Agent Evaluator
Backend Engineer: AI Coding Agent Evaluator

Obsidian • New York (NY)

On-site
USD 454,608 - 647,472
Senior ML Research Engineer - Coding-Agent Benchmarking
Senior ML Research Engineer - Coding-Agent Benchmarking

24-MAG • United States

Remote
USD 600,000 - 1,300,000
Frontier ML Engineer: AI Coding Agent Evaluator
Frontier ML Engineer: AI Coding Agent Evaluator

Mercor • New York (NY)

Hybrid
USD 275,520 - 826,560
AI Code-Agent Evaluator: Frontier DevOps Engineer
AI Code-Agent Evaluator: Frontier DevOps Engineer

Mercor • New York (NY)

On-site
USD 455,000 - 647,000
AI Systems Engineer: Frontier Model Evaluator
AI Systems Engineer: Frontier Model Evaluator

Obsidian • New York (NY)

On-site
USD 180,000 - 220,000
Member of Technical Staff, Coding Research
Member of Technical Staff, Coding Research

Pro Integrate LLC • New York (NY)

Remote
USD 200,000 - 260,000
Equity compensation
Bonuses
Health insurance premium reimbursement
+3
Frontier Coding Data Engineer — Evaluation & Training
Frontier Coding Data Engineer — Evaluation & Training

Surge AI • United States

Remote
USD 140,000 - 210,000