AI Evaluations Engineer — Benchmarking Frontiers

Meta

Menlo Park (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Meta is seeking Research Engineers to join the Evaluations team within Meta Superintelligence Labs. You will curate and build benchmarks for our advanced AI models across text, vision, audio, and beyond, collaborating with world-class researchers to deploy novel evaluation environments.

This is a highly technical role requiring practical research engineering skills and independence. You will design and implement scalable evaluation pipelines, source data, and contribute to tooling that measures

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience.
  • 4+ years of experience in machine learning engineering, machine learning research, or related technical role.
  • Proficiency in Python and experience with ML frameworks such as PyTorch.
  • Experience identifying, designing and completing medium to large technical features independently, without guidance.
  • Demonstrated experience in software engineering practices including version control, testing, and code review practices.
  • Publications at peer-reviewed venues (NeurIPS, ICML, ICLR, ACL, EMNLP, or similar) related to language model evaluation, benchmarking, or deep learning.
  • Hands-on experience with language model post-training and deep learning systems, or building reinforcement learning environments.
  • Experience implementing or developing evaluation benchmarks for large language models and multimodal models (e.g., vision-language, audio, video).
  • Experience working with large-scale distributed systems and data pipelines.
  • Familiarity with language model evaluation frameworks and metrics.
  • Track record of open-source contributions to ML evaluation tools or benchmarks.

Responsibilities

  • Curate and integrate publicly available and internal benchmarks to direct the capabilities of frontier model development.
  • Develop and implement evaluation environments, including environments for novel model capabilities and modalities.
  • Collaborate with external data vendors to source and prepare high-quality evaluation datasets.
  • Execute on the technical vision of research scientists designing new benchmarks and evaluations.
  • Build robust, reusable evaluation pipelines that scale across multiple model lines and product areas.
  • Contribute to evaluation tooling that measures the quality and reliability of evaluation suites.

Skills

Python
ML frameworks
Independent work
Software engineering practices
Publications in peer venues
Reinforcement learning environments
Benchmark development
Multimodal models
Distributed data pipelines

Education

Bachelor's degree in Computer Science/Engineering or equivalent

Tools

PyTorch

Job description

Meta is seeking Research Engineers to join the Evaluations team within Meta Superintelligence Labs. You will curate and build benchmarks for our advanced AI models across text, vision, audio, and beyond, collaborating with world-class researchers to deploy novel evaluation environments.

This is a highly technical role requiring practical research engineering skills and independence. You will design and implement scalable evaluation pipelines, source data, and contribute to tooling that measures

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer - Meta Superintelligence Labs
Research Engineer - Meta Superintelligence Labs

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Safety Evaluation Engineer: Turn AI Risk into Metrics
Safety Evaluation Engineer: Turn AI Risk into Metrics

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Bonus potential
Equity
Benefits
Systems & ML Infra Engineer for Frontier AI Evaluations
Systems & ML Infra Engineer for Frontier AI Evaluations

Meta • Menlo Park (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Staff Engineer - AI Evaluation & Metrics Platform
Staff Engineer - AI Evaluation & Metrics Platform

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000
Research Engineer, Privacy Evals — Frontier AI Safety
Research Engineer, Privacy Evals — Frontier AI Safety

Meta Careers • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Evaluation Engineer: AI Coding Benchmarks & Tests
Evaluation Engineer: AI Coding Benchmarks & Tests

Mercor • United States

Remote
USD 120,000 - 180,000
AI Software Engineer - Evaluation & Benchmarks
AI Software Engineer - Evaluation & Benchmarks

Hire Feed • San Francisco (CA)

On-site
USD 120,000 - 190,000
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
Benchmark Architect for AI Evaluation and Research
Benchmark Architect for AI Evaluation and Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
Benchmarking Research Engineer: Frontier Model Evaluations
Benchmarking Research Engineer: Frontier Model Evaluations

Refresh AI • San Francisco (CA)

On-site
USD 120,000 - 150,000