Senior Coding & AI Model Evaluation Engineer

Braintrust

United States

Remote

USD 103,000 - 124,000

Part time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Braintrust is seeking experienced software engineers for a contracting engagement focused on evaluating and annotating AI models. You will design coding tasks, assess model outputs for code-related tasks, and document failures with rigorous feedback.

This is a senior, hands-on role requiring strong engineering judgment and language fluency. The role centers on real-world software engineering, model evaluation, and applied AI, with a focus on high-quality benchmarking and RL workflows.

Qualifications

  • 5+ years of professional software development experience.
  • Strong Python skills (required).
  • Knowledge of at least one additional programming language.
  • Experience with professional code review, coding annotation, LLM/code evaluation, or benchmark design is a plus, but not required.
  • Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing security patches.
  • Proven ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluent in English (written and spoken).
  • Team lead or mentoring experience is a strong plus.

Responsibilities

  • Evaluate coding tasks involving software vulnerabilities, exploit verification, and security patches.
  • Create high-quality coding prompts and reference answers (benchmark-style, e.g. SWE-Bench-like problems).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and reinforcement learning (RL).
  • Follow detailed annotation and evaluation guidelines with high consistency.

Skills

5+ years software development
Python
Other language
Code review / evaluation
Vulnerability research
English fluency
Team leadership

Job description

Braintrust is seeking experienced software engineers for a contracting engagement focused on evaluating and annotating AI models. You will design coding tasks, assess model outputs for code-related tasks, and document failures with rigorous feedback.

This is a senior, hands-on role requiring strong engineering judgment and language fluency. The role centers on real-world software engineering, model evaluation, and applied AI, with a focus on high-quality benchmarking and RL workflows.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Code Evaluator & Task Designer (Remote)
Senior AI Code Evaluator & Task Designer (Remote)

Braintrust • United States

Remote
USD 96,000 - 117,000
Senior AI Code Evaluator (Python/TypeScript)
Senior AI Code Evaluator (Python/TypeScript)

Turing • United States

Remote
USD 83,000 - 152,000
AI Model Engineer: Code Review & Evaluation
AI Model Engineer: Code Review & Evaluation

AfterQuery Experts • San Francisco (CA)

On-site
USD 110,000 - 170,000
Competitive Pay
AI Software Engineer — Model Evaluation & Code Review
AI Software Engineer — Model Evaluation & Code Review

AfterQuery Experts • New York (NY)

On-site
USD 90,000 - 130,000
Competitive Pay
Remote AI Software Engineer - Code & Data Evaluation
Remote AI Software Engineer - Code & Data Evaluation

Turing • United States

Remote
USD 83,000 - 138,000
Senior AI Coding Evaluation Engineer — Remote Contract
Senior AI Coding Evaluation Engineer — Remote Contract

G2i • United States

Remote
USD 138,000 - 276,000
Senior AI Coding Evaluator - Remote Code Review Pro
Senior AI Coding Evaluator - Remote Code Review Pro

AuraOne • United States

Remote
USD 83,000 - 124,000
AI Software Engineer — Code Review & Model Evaluation
AI Software Engineer — Code Review & Model Evaluation

AfterQuery Experts • San Diego (CA)

On-site
USD 100,000 - 140,000
Competitive Pay
Remote AI Code Evaluation Engineer
Remote AI Code Evaluation Engineer

RemoExperts • United States

Remote
USD 90,000 - 145,000
Remote Senior Software Engineer: AI Code Evaluation
Remote Senior Software Engineer: AI Code Evaluation

24-Mag Llc • United States

Remote
USD 14,000 - 55,000
Fully remote
Flexible hours
Contract-based