Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Braintrust is seeking experienced software engineers for a contracting engagement focused on evaluating and annotating AI models. You will design coding tasks, assess model outputs for code-related tasks, and document failures with rigorous feedback.
This is a senior, hands-on role requiring strong engineering judgment and language fluency. The role centers on real-world software engineering, model evaluation, and applied AI, with a focus on high-quality benchmarking and RL workflows.
This is a contracting engagement - initially 6 months - with potential for long term engagement.
Location: Paris or London-based preferred; alternatively Europe remote for strong candidates
We are building and evaluating state-of-the-art large language models (LLMs) and are looking for experienced software engineers to join our evaluation and annotation team. This role sits at the intersection of real-world software engineering, model evaluation, and applied AI , and is critical to improving model reliability, reasoning, and code quality.
You will design challenging coding tasks, evaluate model outputs against rigorous benchmarks, identify failure modes, and contribute to reinforcement learning and model improvement workflows.
This is not a junior annotation role. We are looking for practitioners with deep hands-on coding experience who can think like both an engineer and an evaluator.
Company: Leading AI Lab