Coding Research Evaluation Engineer

Remotebridge

United States

Remote

USD 200,000 - 260,000

Full time

43 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity compensation
Bonus eligibility
Health-insurance reimbursement
401(k) with company match
Remote-first workforce benefits

Job summary

micro1 seeks a Member of Technical Staff to advance evaluation and development of frontier coding agents. This remote role sits at the intersection of AI research, software engineering, and model evaluation, and you will design benchmarks, methodologies, and data systems that shape how coding models are measured and improved.

You will build evaluation frameworks, lead research initiatives, develop datasets and protocols, analyze model behavior, and create tooling for large-scale experiments.

Qualifications

  • Strong software engineering background with expertise in Python, C++, or comparable programming languages.
  • 3+ years of experience in software engineering, machine learning, AI research, evaluation, or related technical disciplines.
  • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
  • Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
  • Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
  • Strong analytical skills with the ability to investigate model behavior and derive insights from complex technical systems.
  • Excellent written and verbal communication skills, including the ability to clearly articulate technical findings to diverse audiences.
  • Comfortable operating in fast-moving research environments with significant ambiguity and evolving priorities.

Responsibilities

  • Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methodologies, rubrics, and quality standards.
  • Lead end-to-end research initiatives focused on measuring and improving coding model performance across diverse software engineering tasks.
  • Develop high-quality datasets, golden examples, and evaluation protocols that enable reliable assessment of frontier coding systems.
  • Analyze model behavior and failure modes, identifying systematic weaknesses and translating findings into actionable improvements for training and evaluation.
  • Build tooling and infrastructure that support large-scale experimentation, data generation, review workflows, and evaluation pipelines.
  • Establish best practices for coding-agent assessment, ensuring methodological rigor, reproducibility, and measurement quality.
  • Partner closely with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities.
  • Contribute to technical reports, benchmark studies, and client-facing research initiatives that communicate model performance and insights.

Skills

Python
C++
Software engineering
Machine learning
Research

Job description

micro1 seeks a Member of Technical Staff to advance evaluation and development of frontier coding agents. This remote role sits at the intersection of AI research, software engineering, and model evaluation, and you will design benchmarks, methodologies, and data systems that shape how coding models are measured and improved.

You will build evaluation frameworks, lead research initiatives, develop datasets and protocols, analyze model behavior, and create tooling for large-scale experiments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Coding Research
Member of Technical Staff, Coding Research

Remotebridge • United States

Remote
USD 200,000 - 260,000
Equity compensation
Bonus eligibility
Health-insurance reimbursement
+2
Member of Technical Staff, Coding Research
Member of Technical Staff, Coding Research

Pro Integrate LLC • New York (NY)

Remote
USD 200,000 - 260,000
Equity compensation
Bonuses
Health insurance premium reimbursement
+3
Remote | Member of Technical Staff, Coding Research — $600,000–$1,300,000/year
Remote | Member of Technical Staff, Coding Research — $600,000–$1,300,000/year

24-MAG • United States

Remote
USD 600,000 - 1,300,000
Remote Coding Research Engineer - Frontier AI Benchmarks
Remote Coding Research Engineer - Frontier AI Benchmarks

Pro Integrate LLC • New York (NY)

Remote
USD 200,000 - 260,000
Equity compensation
Bonuses
Health insurance premium reimbursement
+3
Senior ML Research Engineer - Coding-Agent Benchmarking
Senior ML Research Engineer - Coding-Agent Benchmarking

24-MAG • United States

Remote
USD 600,000 - 1,300,000
Remote Research Engineer for AI Code & Model Evaluation
Remote Research Engineer for AI Code & Model Evaluation

24-MAG • United States

Remote
USD 69,000 - 138,000
Remote work
Flexible hours
Contractor engagement
Remote Senior Software Engineer: AI Code Evaluation
Remote Senior Software Engineer: AI Code Evaluation

24-Mag Llc • United States

Remote
USD 14,000 - 55,000
Fully remote
Flexible hours
Contract-based
Evaluation Engineer: AI Coding Benchmarks & Tests
Evaluation Engineer: AI Coding Benchmarks & Tests

Mercor • United States

Remote
USD 120,000 - 180,000
Remote Research Engineer - Code Generation & Model Evaluation
Remote Research Engineer - Code Generation & Model Evaluation

RemoteJobsOne • Columbus (OH)

Remote
USD 69,000 - 138,000
Fully remote
Remote Research Engineer: Code Gen & Model Evaluation
Remote Research Engineer: Code Gen & Model Evaluation

Remotebridge • United States

Remote
USD 83,000 - 124,000