Senior AI Code Evaluator & Task Designer (Remote)

Braintrust

United States

Remote

USD 96,000 - 117,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Leading AI Lab is seeking experienced software engineers to join an evaluation and annotation team focused on real-world software engineering, model evaluation, and AI-driven improvements.

Ideal candidates have deep hands-on coding experience, strong Python skills, and a track record in structured evaluation and high-quality feedback. This is a senior, contract-friendly role with collaboration across private LLMs and external models.

Qualifications

  • 5+ years of professional software development experience.
  • Strong Python skills.
  • Knowledge of at least one additional programming language.
  • 1+ year of coding annotation and/or LLM evaluation experience.
  • Prior code reviewer experience is a plus.
  • Fluent in English (written and spoken).
  • Team lead or mentoring experience is a strong plus.

Responsibilities

  • Create high-quality coding prompts and reference answers (benchmark-style, e.g. SWE-Bench-like problems).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and reinforcement learning (RL).
  • Follow detailed annotation and evaluation guidelines with high consistency.

Skills

Python
Code review
Team leadership

Job description

Leading AI Lab is seeking experienced software engineers to join an evaluation and annotation team focused on real-world software engineering, model evaluation, and AI-driven improvements.

Ideal candidates have deep hands-on coding experience, strong Python skills, and a track record in structured evaluation and high-quality feedback. This is a senior, contract-friendly role with collaboration across private LLMs and external models.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer – LLM Evaluation (Fully Remote)
Senior Software Engineer – LLM Evaluation (Fully Remote)

Partner Company • United States

Remote
USD 83,000 - 165,000
Remote-friendly
Flexible hours
Contractor-friendly
Senior Coding & AI Model Evaluation Engineer
Senior Coding & AI Model Evaluation Engineer

Braintrust • United States

Remote
USD 103,000 - 124,000
Senior AI Coding Evaluation Engineer — Remote Contract
Senior AI Coding Evaluation Engineer — Remote Contract

G2i • United States

Remote
USD 138,000 - 276,000
Remote Senior Software Engineer: AI Code Evaluation
Remote Senior Software Engineer: AI Code Evaluation

24-Mag Llc • United States

Remote
USD 14,000 - 55,000
Fully remote
Flexible hours
Contract-based
Senior Coding Annotator / LLM Evaluation Engineer (Contract / BYO)
Senior Coding Annotator / LLM Evaluation Engineer (Contract / BYO)

Braintrust • United States

Remote
USD 96,000 - 117,000
AI Evaluator - Python (Freelance Opportunity)
AI Evaluator - Python (Freelance Opportunity)

Biz Tech Consultants • United States

Remote
USD 55,000 - 110,000
Remote AI Software Engineer - Code & Data Evaluation
Remote AI Software Engineer - Code & Data Evaluation

Turing • United States

Remote
USD 83,000 - 138,000
AI Evaluator - Freelance Opportunity (Python)
AI Evaluator - Freelance Opportunity (Python)

Biz Tech Consultants • United States

Remote
USD 120,000 - 180,000
Senior AI Code Interaction Evaluator - Remote
Senior AI Code Interaction Evaluator - Remote

G2i • United States

Remote
AUD 198,000 - 396,000
AI Evaluator - JavaScript (Freelance Opportunity)
AI Evaluator - JavaScript (Freelance Opportunity)

Biz Tech Consultants • United States

Remote
USD 83,000 - 152,000