Principal Coding Annotator / LLM Evaluation Engineer

Braintrust

United States

Remote

USD 103,000 - 124,000

Part time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Braintrust is seeking experienced software engineers for a contracting engagement focused on evaluating and annotating AI models. You will design coding tasks, assess model outputs for code-related tasks, and document failures with rigorous feedback.

This is a senior, hands-on role requiring strong engineering judgment and language fluency. The role centers on real-world software engineering, model evaluation, and applied AI, with a focus on high-quality benchmarking and RL workflows.

Qualifications

  • 5+ years of professional software development experience.
  • Strong Python skills (required).
  • Knowledge of at least one additional programming language.
  • Experience with professional code review, coding annotation, LLM/code evaluation, or benchmark design is a plus, but not required.
  • Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing security patches.
  • Proven ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluent in English (written and spoken).
  • Team lead or mentoring experience is a strong plus.

Responsibilities

  • Evaluate coding tasks involving software vulnerabilities, exploit verification, and security patches.
  • Create high-quality coding prompts and reference answers (benchmark-style, e.g. SWE-Bench-like problems).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and reinforcement learning (RL).
  • Follow detailed annotation and evaluation guidelines with high consistency.

Skills

5+ years software development
Python
Other language
Code review / evaluation
Vulnerability research
English fluency
Team leadership

Job description

  • Rate: $75 – $90/hr
  • Hours: 20 hours / week
  • Experience: 10+ years
  • Location: Belgium | Denmark + 10 more

This is a contracting engagement - initially 6 months - with potential for long term engagement.

Location: Paris or London-based preferred; alternatively Europe remote for strong candidates

We are building and evaluating state-of-the-art large language models (LLMs) and are looking for experienced software engineers to join our evaluation and annotation team. This role sits at the intersection of real-world software engineering, model evaluation, and applied AI , and is critical to improving model reliability, reasoning, and code quality.

You will design challenging coding tasks, evaluate model outputs against rigorous benchmarks, identify failure modes, and contribute to reinforcement learning and model improvement workflows.

This is not a junior annotation role. We are looking for practitioners with deep hands-on coding experience who can think like both an engineer and an evaluator.

What You’ll Do
  • Evaluate coding tasks involving software vulnerabilities, exploit verification, and security patches.
  • Create high-quality coding prompts and reference answers (benchmark-style, e.g. SWE-Bench-like problems).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and reinforcement learning (RL).
  • Follow detailed annotation and evaluation guidelines with high consistency.
What We’re Looking For
  • 5+ years of professional software development experience.
  • Strong Python skills (required).
  • Knowledge of at least one additional programming language (bonus).
  • Experience with professional code review, coding annotation, LLM/code evaluation, or benchmark design is a plus, but not required.
  • Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting, or validating security patches.
  • Proven ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluent in English (written and spoken).
  • Team lead or mentoring experience is a strong plus.
Why This Role
  • Work hands-on with cutting-edge LLMs.
  • Apply real-world engineering judgment to model evaluation and improvement.
  • High-impact, technical work with a focused, senior team.

Company: Leading AI Lab

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Coding Annotator / LLM Evaluation Engineer (Contract / BYO)
Senior Coding Annotator / LLM Evaluation Engineer (Contract / BYO)

Braintrust • United States

Remote
USD 96,000 - 117,000
Embedded / Systems Engineer (C) – AI Code Analysis | Remote
Embedded / Systems Engineer (C) – AI Code Analysis | Remote

Crossing Hurdles • United States

On-site
USD 82,656 - 137,760
Senior AI Code Evaluator & Task Designer (Remote)
Senior AI Code Evaluator & Task Designer (Remote)

Braintrust • United States

Remote
USD 96,000 - 117,000
Senior Application Security Engineer - AI Code Evaluation
Senior Application Security Engineer - AI Code Evaluation

Braintrust • United States

Remote
USD 103,000 - 124,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • New York (NY)

On-site
USD 68,880 - 206,640
Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

24-Mag Llc • United States

Remote
USD 14,000 - 55,000
Fully remote
Flexible hours
Contract-based
Senior Software Engineer – LLM Evaluation (Fully Remote)
Senior Software Engineer – LLM Evaluation (Fully Remote)

Partner Company • United States

Remote
USD 83,000 - 165,000
Remote-friendly
Flexible hours
Contractor-friendly
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • San Francisco (CA)

On-site
USD 68,880 - 206,640
Remote | Senior Software Engineer – LLM Evaluation
Remote | Senior Software Engineer – LLM Evaluation

24-Mag Llc • New York (NY)

Remote
USD 14,000 - 55,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • New York (NY)

On-site
USD 68,880 - 206,640