AI Safety Benchmark Engineer (Evals)

White Circle

Paris

Hybride

EUR 90 000 - 130 000

Plein temps

14 jours+
Générateur de candidature

N’envoyez pas de CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Avantages offerts par ce poste

Equity
Flexible time off
Hybrid office Paris/London
Relocation support
Health insurance
Mental health support
Office meals

Résumé du poste

White Circle, an AI Safety company, seeks a Research Engineer to own our internal benchmarks for single/multi-turn content guardrails and agent safety. You will build evals for flagship models and extend benchmarks as product data evolves, collaborating with the product team and research groups.

You will craft synthetic data, reproduce published benchmarks, and maintain production-grade code with clean abstractions. Office in Paris with hybrid setup; relocation support may apply.

Qualifications

  • Built an LLM benchmark from scratch that distinguished model capabilities
  • Generated synthetic data for post-training textual or multimodal models
  • Could reproduce a published benchmark result and assess the methodology's robustness
  • Proficient in Python with production-grade code and clean abstractions

Responsabilités

  • Own and maintain internal benchmark suite for single/multi-turn content guardrails and agentic safety
  • Build benchmarks that distinguish specific model capabilities
  • Collaborate with product team to build evals for flagship models
  • Extend evals to new features and data across verticals
  • Study and quantify realistic agentic and LLM failure modes in the wild

Connaissances

Python
LLM Benchmarking
Production-ready code
Parallel inference
Agent safety
Frontier models
Research experience
Data synthesis

Outils

Git
Docker
Kubernetes
PyTorch

Description du poste

White Circle, an AI Safety company, seeks a Research Engineer to own our internal benchmarks for single/multi-turn content guardrails and agent safety. You will build evals for flagship models and extend benchmarks as product data evolves, collaborating with the product team and research groups.

You will craft synthetic data, reproduce published benchmarks, and maintain production-grade code with clean abstractions. Office in Paris with hybrid setup; relocation support may apply.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Benchmark Engineer, AI Safety & LLM Evaluation
Benchmark Engineer, AI Safety & LLM Evaluation

Aisafety • Paris

Hybride
EUR 80 000 - 110 000
Relocation package
Hybrid work (Paris)
Comprehensive medical insurance
+1
Research Engineer (Evals)
Research Engineer (Evals)

Aisafety • Paris

Sur place
EUR 80 000 - 110 000
Relocation package
Hybrid work (Paris)
Comprehensive medical insurance
+1
Research Engineer (Evals)
Research Engineer (Evals)

White Circle • Paris

Hybride
EUR 90 000 - 130 000
Equity
Flexible time off
Hybrid office Paris/London
+4
Research Scientist, AI Safety & Multi-Agent Environments
Research Scientist, AI Safety & Multi-Agent Environments

White Circle • Paris

Hybride
EUR 131 371 - 218 952
Paid time off
Comprehensive medical insurance
Team off-sites twice a year
Research Scientist: Frontier AI Risk & Evaluations
Research Scientist: Frontier AI Risk & Evaluations

Safer Ai • Paris

Hybride
EUR 90 000 - 140 000
Health insurance
Retirement plans
50% home transportation
+2
QA Engineer for AI Safety — Hybrid Paris/London, Equity
QA Engineer for AI Safety — Hybrid Paris/London, Equity

Npv • Paris

Hybride
EUR 60 000 - 90 000
Equity
Hybrid work
Relocation package
+3
QA Engineer – AI Safety Testing (Playwright, Python) | Hybrid Paris
QA Engineer – AI Safety Testing (Playwright, Python) | Hybrid Paris

White Circle • Paris

Hybride
EUR 42 944 - 68 711
Paid time off according to local regulations
Relocation package
Best medical insurance in France
+2
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • Paris

Sur place
EUR 70 000 - 90 000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • Paris

Sur place
EUR 70 000 - 110 000
Backend Rust Engineer for AI Safety at Scale
Backend Rust Engineer for AI Safety at Scale

Whitecircle • Paris

Hybride
EUR 90 000 - 130 000
Equity package
Relocation support
Hybrid work from Paris