Senior LLM Evaluation Engineer & Coding Specialist

Braintrust

Netherlands

On-site

EUR 90,000 - 150,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Braintrust is seeking seasoned software engineers for a contracting role focused on evaluation and annotation of state-of-the-art LLMs. You will design coding prompts, assess model outputs, and document failures while working at a senior level with a focused, highly capable team.

Ideal candidates have 10+ years in software development, expert Python skills, and hands-on experience with code evaluation or LLM benchmarking.

Qualifications

  • 10+ years of professional software development experience.
  • Strong Python skills (required).
  • Knowledge of at least one additional programming language.
  • 1+ year of coding annotation and/or LLM evaluation experience (part-time OK) for a major frontier AI lab or AI infrastructure company.
  • Prior code reviewer experience is a plus.
  • Proven ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluent in English (written and spoken).
  • Team lead or mentoring experience is a strong plus.

Responsibilities

  • Create high-quality coding prompts and reference answers (benchmark-style, e.g. SWE-Bench-like problems).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and reinforcement learning (RL).
  • Follow detailed annotation and evaluation guidelines with high consistency.

Skills

Python
English fluency
Other programming language

Job description

Braintrust is seeking seasoned software engineers for a contracting role focused on evaluation and annotation of state-of-the-art LLMs. You will design coding prompts, assess model outputs, and document failures while working at a senior level with a focused, highly capable team.

Ideal candidates have 10+ years in software development, expert Python skills, and hands-on experience with code evaluation or LLM benchmarking.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote LLM Evaluation Engineer - AI Research & Benchmarks
Remote LLM Evaluation Engineer - AI Research & Benchmarks

Blue Lynx Employment BV • Netherlands

Remote
EUR 60,000 - 80,000
25 holidays per annum
Pension plan
Opportunity to work on cutting-edge AI initiatives
Software Engineer, LLM Evaluation – English
Software Engineer, LLM Evaluation – English

Blue Lynx Employment BV • Netherlands

Remote
EUR 60,000 - 80,000
25 holidays per annum
Pension plan
Opportunity to work on cutting-edge AI initiatives
Founding AI Engineer: LLM Infra, Evals & Observability
Founding AI Engineer: LLM Infra, Evals & Observability

Jobtailor • Amsterdam

On-site
EUR 120,000 - 180,000
Senior Enterprise AI Architect — LLM, RAG & Agentic Systems
Senior Enterprise AI Architect — LLM, RAG & Agentic Systems

Jobtailor • Zaandam

On-site
EUR 120,000 - 180,000
Staff Research Engineer (LLM Pre-Training)
Staff Research Engineer (LLM Pre-Training)

Jetbrains • Amsterdam

On-site
EUR 70,000 - 90,000
Senior LLM Architect
Senior LLM Architect

IC Resources • Eindhoven

Hybrid
EUR 70,000 - 100,000
Competitive salary DOE
Equity
Hybrid working arrangements
+1
Founding AI Engineer
Founding AI Engineer

Jobtailor • Amsterdam

On-site
EUR 120,000 - 180,000
Remote Legal Expert - 42119
Remote Legal Expert - 42119

Turing • Netherlands

On-site
Senior Agentic AI Engineer: Scalable LLM & RAG Solutions
Senior Agentic AI Engineer: Scalable LLM & RAG Solutions

GeekSoft Consulting • Amsterdam

Hybrid
EUR 110,000 - 170,000
Dutch benefits
Career development
International culture
Senior Data Scientist — LLM/ML Production Lead
Senior Data Scientist — LLM/ML Production Lead

Moss • Amsterdam

On-site
EUR 120,000 - 180,000
Equity
In-person culture
Learning budget