Senior LLM Evaluation Engineer & Coding Annotator

Braintrust

Germany (OH)

On-site

USD 111,032 - 190,341

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Braintrust seeks an experienced software engineer to join our evaluation and annotation team. You will work at the intersection of real-world software engineering, model evaluation, and applied AI, contributing to reliability and code quality.

This contracting engagement is initially six months with potential for long-term engagement. Location preferences are Paris or London, with Europe-wide remote options for highly qualified candidates.

Qualifications

  • 10+ years of professional software development experience.
  • Strong Python skills required.
  • 1+ year of coding annotation and/or LLM evaluation experience (part-time OK).
  • Fluent in English (written and spoken).

Responsibilities

  • Create high-quality coding prompts and reference answers (benchmark-style).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and RL.
  • Follow detailed annotation and evaluation guidelines with high consistency.

Skills

Python
Software development
Code evaluation / LLM evaluation
English fluency
Team leadership
Mentoring
Code review
LLM evaluation tools

Education

Bachelor's degree in CS or related field

Tools

Mistral-based LLMs

Job description

Braintrust seeks an experienced software engineer to join our evaluation and annotation team. You will work at the intersection of real-world software engineering, model evaluation, and applied AI, contributing to reliability and code quality.

This contracting engagement is initially six months with potential for long-term engagement. Location preferences are Paris or London, with Europe-wide remote options for highly qualified candidates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Coding Annotator & LLM Evaluation Engineer - Remote
Senior Coding Annotator & LLM Evaluation Engineer - Remote

Braintrust • Town of Belgium (WI)

On-site
USD 104,000 - 173,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • San Francisco (CA)

Remote
Embedded / Systems Engineer (C) – AI Code Analysis | Remote
Embedded / Systems Engineer (C) – AI Code Analysis | Remote

Crossing Hurdles • United States

Remote
AI Engineer - LLM Training & Evaluation (Remote)
AI Engineer - LLM Training & Evaluation (Remote)

Prolific • Memphis (TN)

Hybrid
USD 90,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home
Senior LLM Evaluation & Fine-Tuning Architect
Senior LLM Evaluation & Fine-Tuning Architect

Innodata Inc. • United States

On-site
USD 140,000 - 160,000
Senior GenAI & LLM Evaluation Engineer - Remote US
Senior GenAI & LLM Evaluation Engineer - Remote US

Acuity Analytics • United States

On-site
USD 140,000 - 190,000
Remote Senior Python Engineer – LLM Evaluation (US-based)
Remote Senior Python Engineer – LLM Evaluation (US-based)

Turing • Chicago (IL)

On-site
Senior LLM Evaluation & Quality Engineer
Senior LLM Evaluation & Quality Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • New York (NY)

Remote
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required