Remote LLM Evaluation Scientist - Benchmark Frontiers

Anyone AI

Paris

Sur place

EUR 90 000 - 140 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

Anyone AI Labs is seeking a dedicated Research Scientist to advance how frontier models are evaluated. You will design benchmarks, validate methodologies, and build evaluation packages across reasoning, coding, agents, and multi-modal systems.

You will lead expert recruitment, coordinate with labs and CEOs, and push for public benchmarks and papers at venues like NeurIPS and ICLR. Fluent English required; Spanish is a plus.

Qualifications

  • Experience in ML evaluation or benchmarking research, with a track record of open benchmarks or eval studies.
  • Deep expertise in LLM/frontier-model benchmarking, including code-model and agentic evaluation.
  • Fluency with measurement concepts: construct validity, psychometrics, rubrics, headroom, contamination.
  • Interest in safety-oriented evaluation, robustness, and measuring hard-to-measure capabilities.
  • Ability to lead a team or expert pool to high standards; strong written and oral communication.

Responsabilités

  • Evaluation research: define targets, design benchmarks, ensure validity and reliability.
  • Benchmark development: create packages with expert-verified ground truth and multi-model headroom.
  • Recruit, calibrate, and review a pool of experts in coding, agents, and STEM/reasoning.
  • Serve as technical point of contact for labs, with CEO support, translating needs into designs.
  • Deliver and disseminate: pilot projects, public benchmarks, and papers for venues like NeurIPS, ICLR, ACL.

Connaissances

ML evaluation
Benchmarking
LLM evaluation
Scientific research
English fluency
Spanish language

Description du poste

Anyone AI Labs is seeking a dedicated Research Scientist to advance how frontier models are evaluated. You will design benchmarks, validate methodologies, and build evaluation packages across reasoning, coding, agents, and multi-modal systems.

You will lead expert recruitment, coordinate with labs and CEOs, and push for public benchmarks and papers at venues like NeurIPS and ICLR. Fluent English required; Spanish is a plus.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Remote AI Benchmark Engineer & Researcher
Remote AI Benchmark Engineer & Researcher

Pathway • Paris

À distance
EUR 75 000 - 115 000
Intellectually stimulating work environment
Flexible remote work setup
Compensation based on profile and location
Senior AI SME & Localization Engineer
Senior AI SME & Localization Engineer

LILT AI • Paris

Sur place
EUR 50 000 - 70 000
AI Research Scientist (Paris)
AI Research Scientist (Paris)

Lexsi Labs • Paris

Sur place
EUR 60 000 - 90 000
AI Research Scientist (Paris)
AI Research Scientist (Paris)

Lexsi Labs • Paris

Sur place
EUR 85 000 - 130 000
Technical Staff Member – Agent
Technical Staff Member – Agent

Jobtailor • Paris

Sur place
EUR 120 000 - 180 000
Subject Matter Expert – Quantative/Scientific/Corporate (English/French) – Remote
Subject Matter Expert – Quantative/Scientific/Corporate (English/French) – Remote

LILT AI • Paris

Sur place
EUR 50 000 - 70 000
Research Engineer - FAIR, SGT
Research Engineer - FAIR, SGT

Meta • Paris

Sur place
EUR 90 000 - 130 000
Robotics SME (French) - AI Benchmarking Expert
Robotics SME (French) - AI Benchmarking Expert

LILT AI • Paris

Sur place
EUR 60 000 - 80 000
Remote AI Evaluation Specialist - Rubric Expert & Insights
Remote AI Evaluation Specialist - Rubric Expert & Insights

Crossing Hurdles • La Réunion

Sur place
EUR 42 000 - 64 000
Research Engineer (Evals)
Research Engineer (Evals)

Aisafety • Paris

Hybride
EUR 80 000 - 110 000
Relocation package
Hybrid work (Paris)
Comprehensive medical insurance
+1