Senior LLM Evaluation & Coding Engineer (Remote EU)

Braintrust

United Kingdom

On-site

GBP 70,706 - 141,412

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Braintrust is hiring experienced software engineers to join our evaluation and annotation team for a contracting engagement. The role focuses on real-world software engineering, model evaluation, and applied AI, aiming to improve model reliability, reasoning, and code quality.

You will design challenging coding tasks, evaluate model outputs against rigorous benchmarks, identify failure modes, and contribute to reinforcement learning and model improvement workflows.

Qualifications

  • 10+ years of professional software development experience.
  • Strong Python skills are required.
  • At least 1 year of coding annotation and/or LLM evaluation experience.
  • Ability to lead or mentor others is a plus.
  • Fluent in English (written and spoken).

Responsibilities

  • Create high-quality coding prompts and reference answers (benchmark-style, e.g. SWE-Bench-like problems).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and reinforcement learning (RL).
  • Follow detailed annotation and evaluation guidelines with high consistency.

Skills

Python
Code review
LLM evaluation
English fluency
Team leadership

Tools

Git
Python tooling

Job description

Braintrust is hiring experienced software engineers to join our evaluation and annotation team for a contracting engagement. The role focuses on real-world software engineering, model evaluation, and applied AI, aiming to improve model reliability, reasoning, and code quality.

You will design challenging coding tasks, evaluate model outputs against rigorous benchmarks, identify failure modes, and contribute to reinforcement learning and model improvement workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Evaluation Engineer - Code & AI Systems
Senior LLM Evaluation Engineer - Code & AI Systems

Braintrust • Greater London

On-site
GBP 90,000 - 150,000
Remote AI Software Engineer - LLM Data & Code Curation
Remote AI Software Engineer - LLM Data & Code Curation

Turing • Greater London

Remote
GBP 41,000 - 81,000
Remote Senior Software Engineer: AI Code Evaluation Systems
Remote Senior Software Engineer: AI Code Evaluation Systems

Synthires • Greater London

On-site
GBP 102,000 - 153,000
Fully remote work
Weekly payments
Remote Senior ML Expert - LLM Reasoning & Training Data
Remote Senior ML Expert - LLM Reasoning & Training Data

Alignerr • Cambridge

On-site
GBP 83,000 - 179,000
Staff, Data Quality & LLM Evaluation (Remote-friendly)
Staff, Data Quality & LLM Evaluation (Remote-friendly)

Cohere • Greater London

On-site
GBP 60,000 - 90,000
Remote-friendly work environment
Daily lunch program for office workers
Co-working benefit for remote employees
Principal AI Engineer: LLM Tutor Systems (Remote)
Principal AI Engineer: LLM Tutor Systems (Remote)

RecT Solutions • Greater London

On-site
GBP 120,000 - 190,000
Senior AI Engineer (LLMs) - Remote & Impactful
Senior AI Engineer (LLMs) - Remote & Impactful

United States Digital Space LLC • United Kingdom

Hybrid
GBP 70,000 - 110,000
Remote work flexibility
Competitive salary
Comprehensive benefits package
+1
English LLM Data Annotator (Remote, Entry level OK, Freelance)
English LLM Data Annotator (Remote, Entry level OK, Freelance)

T-Maxx International • United Kingdom

Remote
GBP 17,000 - 30,000
Flexible hours
Entry-level path into tech
Training and guidance
Senior AI Software Engineer - Python, LLM Evaluation & Recs
Senior AI Software Engineer - Python, LLM Evaluation & Recs

Jobtailor • Greater London

On-site
GBP 80,000 - 110,000
AI/ML Engineer — RLHF & Model Evaluation (Remote)
AI/ML Engineer — RLHF & Model Evaluation (Remote)

Prolific • Greater London

On-site
GBP 80,000 - 100,000