LLM Evaluation Scientist: Frontier Benchmarking

Anyone AI

Buenos Aires

Presencial

ARS 178.909.546 - 268.364.320

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Anyone AI Labs is seeking a Research Scientist to lead evaluations and benchmarking for large language models. You will design frontier-grade evaluation methods and benchmarks across reasoning, coding, agents, tool use, and multimodal tasks, grounded in expert-verified truth and validated across multiple models.

You will own the study lifecycle from framing questions to publishing results, with opportunities to contribute to public benchmarks and papers at top venues, while collaborating with

Formación

  • Research background in ML evaluation or benchmarking.
  • Deep LLM/frontier-model benchmarking expertise, especially code-model and agentic evaluation.
  • Fluency with measurement concepts: construct validity, rubrics, headroom, contamination.

Responsabilidades

  • Evaluation research: design new benchmark targets and evaluation methods.
  • Benchmark development: build evaluation packages with ground truth and multi-model results.
  • Experts: recruit and review a pool across coding, agents, and reasoning.
  • Lab relationships: act as technical contact for labs with CEO support.
  • Delivery & dissemination: translate lab requests into pilots and publish benchmarks/papers.

Conocimientos

ML evaluation
Benchmarking
English fluency

Descripción del empleo

Anyone AI Labs is seeking a Research Scientist to lead evaluations and benchmarking for large language models. You will design frontier-grade evaluation methods and benchmarks across reasoning, coding, agents, tool use, and multimodal tasks, grounded in expert-verified truth and validated across multiple models.

You will own the study lifecycle from framing questions to publishing results, with opportunities to contribute to public benchmarks and papers at top venues, while collaborating with

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Native-Language AI Benchmark Engineer
Remote Native-Language AI Benchmark Engineer

Lilt • Argentina

A distancia
ARS 124.818.000 - 249.637.000
Remote work
Flexible schedule
Prompt payments
Remote AI Quality Evaluator & Support for LLM Benchmarking
Remote AI Quality Evaluator & Support for LLM Benchmarking

Innodata Inc. • Argentina

Presencial
ARS 29.181.000 - 52.109.000
Senior NLP Scientist: Production-Grade LLMs
Senior NLP Scientist: Production-Grade LLMs

Sourceability • Municipio de Esquel

Presencial
ARS 176.967.000 - 353.936.000
Lead AI Engineer: Scalable LLM & Agentic Systems
Lead AI Engineer: Scalable LLM & Agentic Systems

EPAM Systems • Argentina

Presencial
ARS 3.000.000 - 5.000.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Lead AI Engineer: Scalable LLM & AI Systems
Lead AI Engineer: Scalable LLM & AI Systems

EPAM Systems • Argentina

Presencial
ARS 1.000.000 - 2.000.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Senior NLP Scientist: Production-Ready LLMs & AI
Senior NLP Scientist: Production-Ready LLMs & AI

Sourceability • Buenos Aires

Presencial
ARS 3.000.000 - 6.000.000
Senior AI Engineer: Architect Scalable LLM Solutions
Senior AI Engineer: Architect Scalable LLM Solutions

EPAM Systems, Inc. • Argentina

A distancia
ARS 1.800.000 - 3.200.000
Senior AI/ML Engineer - Build LLM Agents & Guardrails
Senior AI/ML Engineer - Build LLM Agents & Guardrails

Ciklum • Argentina

Presencial
ARS 3.000.000 - 6.000.000
Medical insurance
Mental health programs
Udemy license
+1
Senior AI Engineer - Agentic LLMs & RAG in Production
Senior AI Engineer - Agentic LLMs & RAG in Production

intive • Argentina

Presencial
ARS 1.200.000 - 2.400.000
Remote-friendly work environment
Udemy access & training resources
Mentorship program
+1
AI Engineer — Vertex AI, LLMs & Python
AI Engineer — Vertex AI, LLMs & Python

Qodea • Buenos Aires

Presencial
ARS 97.269.000 - 142.162.000
OSDE 210 for family group
Work from Home Allowance
Birthday leave
+2