LLM Evaluation & Benchmarking Scientist

Anyone AI

Colombia

Presencial

COP 367.462.836 - 567.897.110

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Remote work
Opportunity to publish

Descripción de la vacante

Anyone AI Labs is seeking a Research Scientist focused on LLM evaluations and benchmarking. You will design frontier-grade evaluation packages across reasoning, coding, agents, and multi-modal capabilities, grounded in expert-verified truth and validated against multiple models.

You’ll own evaluation methodology, collaborate with labs, and contribute to public benchmarks and papers, with a strong emphasis on rigor and QC in a remote, LatAm/US‑centric setup.

Formación

  • Track record of published or open benchmarks, eval/measurement research, or equivalent hands-on work that labs have relied on.
  • Deep LLM/frontier-model benchmarking expertise, with real strength in code‑model and agentic evaluation.
  • Fluency with the measurement problem: construct validity, psychometrics, rubrics, headroom, contamination, and what makes a task discriminate a model.
  • Interest in the safety side of evaluation, capability elicitation, robustness, and measuring the hardest-to-measure things.
  • Proven ability to hold a team or expert pool to a rigorous standard.
  • Comfort with the full research loop: framing the question, running the study, and writing it up.
  • Fluent English; Spanish a plus.

Responsabilidades

  • Evaluation research: design original benchmark targets and measures.
  • Benchmark development with expert-verified ground truth and multi-model headroom results.
  • Recruit, calibrate, and review a pool of experts across coding, agentic/tool-use, and STEM/reasoning.
  • Act as technical point of contact for labs and translate needs into evaluation designs.
  • Turn lab requests into pilots and publish public benchmarks and papers when applicable.

Conocimientos

ML evaluation research
Frontier-model benchmarking
Code-model evaluation
Agentic evaluation
Measurement theory
Safety in evaluation
Team leadership
English fluency
Spanish (plus)

Descripción del empleo

Anyone AI Labs is seeking a Research Scientist focused on LLM evaluations and benchmarking. You will design frontier-grade evaluation packages across reasoning, coding, agents, and multi-modal capabilities, grounded in expert-verified truth and validated against multiple models.

You’ll own evaluation methodology, collaborate with labs, and contribute to public benchmarks and papers, with a strong emphasis on rigor and QC in a remote, LatAm/US‑centric setup.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Senior AI Engineer - LLM & Agentic AI
Remote Senior AI Engineer - LLM & Agentic AI

Applaudo • Colombia

Presencial
COP 282.841.000 - 408.548.000
Senior AI Engineer: LLMs & Agentic AI Architect (Remote)
Senior AI Engineer: LLMs & Agentic AI Architect (Remote)

Applaudo • Bogotá

Presencial
COP 120.000.000 - 210.000.000
Senior AI Engineer: LLMs & Agentic AI
Senior AI Engineer: LLMs & Agentic AI

Applaudo • Medellín

Presencial
COP 40.000.000 - 65.000.000
Senior AI Engineer: GenAI & LLM Evaluation (Remote)
Senior AI Engineer: GenAI & LLM Evaluation (Remote)

BlackCube Labs • Colombia

Presencial
COP 286.205.000 - 413.407.000
Contrato directo indefinido
Senior AI Engineer: Production LLMs & Agentic AI (Remote)
Senior AI Engineer: Production LLMs & Agentic AI (Remote)

Werben HR • Colombia

A distancia
COP 12.000.000 - 24.000.000
Fully remote work
Cutting-edge AI projects
International distributed teams
Senior LLM Engineer - Remote AI Architect (RAG/MLOps)
Senior LLM Engineer - Remote AI Architect (RAG/MLOps)

BairesDev • Antioquia

A distancia
COP 281.171.000 - 468.618.000
Remote work for global team
USD or local currency compensation
Hardware and software setup for home
Senior AI Engineer (LLM/GenAI)
Senior AI Engineer (LLM/GenAI)

First Line Software • Bogotá

Híbrido
COP 120.000.000 - 210.000.000
Senior LLM Engineer — Remote, Flexible Hours, Global Impact
Senior LLM Engineer — Remote, Flexible Hours, Global Impact

BairesDev • Colombia

Presencial
COP 376.400.000 - 564.600.000
Remote work
USD or local currency
Home office setup
+4
Remote AI Prompt Architect for LLM Precision
Remote AI Prompt Architect for LLM Precision

BairesDev • Antioquia

A distancia
COP 281.171.000 - 468.618.000
100% remote work
USD or local currency compensation
Home office setup provided
+4
Lead AI Engineer - Enterprise LLM Solutions Architect
Lead AI Engineer - Enterprise LLM Solutions Architect

EPAM Systems • Colombia

Presencial
COP 120.000.000 - 180.000.000
Health coverage
Stock option plan
Learning culture
+1