Senior Coding Annotator & LLM Evaluation Engineer - Remote

Braintrust

Town of Belgium (WI)

On-site

USD 104,000 - 173,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Braintrust is seeking an experienced software engineer to join our evaluation and annotation team, focusing on real-world coding tasks and model evaluation. This role requires hands-on coding expertise, the ability to think like an engineer and evaluator, and strong English communication.

The engagement is contracting for 6 months with potential for extension. Location options include Paris/London or Europe remote for top candidates, offering a high-impact, senior team environment focused on

Qualifications

  • 10+ years of professional software development experience.
  • Strong Python skills (required).
  • Knowledge of at least one additional programming language (bonus).
  • 1+ year of coding annotation and/or LLM evaluation experience (part-time OK).
  • Prior code reviewer experience is a plus.
  • Fluent in English (written and spoken).
  • Team lead or mentoring experience is a strong plus.

Responsibilities

  • Create high-quality coding prompts and reference answers (benchmark-style).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and RL.
  • Follow detailed annotation and evaluation guidelines with high consistency.

Skills

Python
LLM evaluation
Code review
Benchmark design
English fluency

Tools

Mistral-based LLMs

Job description

Braintrust is seeking an experienced software engineer to join our evaluation and annotation team, focusing on real-world coding tasks and model evaluation. This role requires hands-on coding expertise, the ability to think like an engineer and evaluator, and strong English communication.

The engagement is contracting for 6 months with potential for extension. Location options include Paris/London or Europe remote for top candidates, offering a high-impact, senior team environment focused on

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Evaluation Engineer & Coding Annotator
Senior LLM Evaluation Engineer & Coding Annotator

Braintrust • Germany (OH)

On-site
Embedded / Systems Engineer (C) – AI Code Analysis | Remote
Embedded / Systems Engineer (C) – AI Code Analysis | Remote

Crossing Hurdles • United States

Remote
Remote Code Annotation Specialist (Part‑Time)
Remote Code Annotation Specialist (Part‑Time)

cohere • United States

On-site
Senior AI Code Evaluation Engineer (Remote)
Senior AI Code Evaluation Engineer (Remote)

YO AI Labs • Town of Texas (WI)

Remote
USD 83,000 - 152,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • San Francisco (CA)

Remote
C++ Systems Engineer – AI Model Evaluation & Code Review | Remote
C++ Systems Engineer – AI Model Evaluation & Code Review | Remote

Crossing Hurdles • United States

On-site
AI Engineer - LLM Training & Evaluation (Remote)
AI Engineer - LLM Training & Evaluation (Remote)

Prolific • Memphis (TN)

Hybrid
USD 90,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home
Remote Senior Python Engineer – LLM Evaluation (US-based)
Remote Senior Python Engineer – LLM Evaluation (US-based)

Turing • Chicago (IL)

On-site
Senior Software Engineer — AI Coding Environments (Remote)
Senior Software Engineer — AI Coding Environments (Remote)

YO AI Labs • California (MO)

Remote
USD 83,000 - 165,000
Remote Senior Software Engineer (Contract) — AI Coding Environments
Remote Senior Software Engineer (Contract) — AI Coding Environments

YO AI Labs • Austin (TX)

Remote
USD 70,000 - 100,000