Research Engineer, Benchmarks

Clera

Singapore

On-site

SGD 190,000 - 317,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Visa sponsorship available

Job summary

Clera seeks a research/ML engineer to design and implement rigorous, domain-specific benchmarks evaluating frontier AI agents on realistic workflows. You will work with a small, highly technical team to deliver evaluations trusted by AI labs and enterprise customers.

The role focuses on building benchmarks, infrastructure, and metrics while collaborating with domain experts to translate complex workflows into evaluation tasks and delivering thorough benchmark reports.

Qualifications

  • 2 to 4 years of experience in research engineering or ML engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.
  • Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.
  • Demonstrated experience designing and running benchmarks or evaluation environments for AI agents or large language models.
  • Experience building infrastructure to reliably run AI models or agents against evaluation tasks at scale.
  • Experience developing metrics or validation studies to assess benchmark difficulty, reliability, and real-world correlation.
  • Ability to collaborate with domain experts and translate complex workflows into evaluation criteria.

Responsibilities

  • Design, implement, and maintain high-quality internal benchmarks for evaluating frontier agents on domain-specific tasks.
  • Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.
  • Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.
  • Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
  • Validate that benchmark performance correlates with real-world evaluations and customer needs.
  • Write clear technical documentation and benchmark reports for research and engineering audiences.

Skills

Research engineering
ML engineering
Benchmark design
Collaboration
Technical documentation

Tools

Python
Docker
Linux

Job description

About the Role

You will own the design and implementation of rigorous, domain-specific benchmarks used to evaluate frontier AI agents on realistic workflows. Sitting within a small, highly technical team of researchers and engineers, this role is central to delivering evaluations that AI labs and enterprise customers genuinely trust and rely on.

What You'll Do
  • Design, implement, and maintain high-quality internal benchmarks for evaluating frontier agents on domain-specific tasks.

  • Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.

  • Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.

  • Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.

  • Validate that benchmark performance correlates with real-world evaluations and customer needs.

  • Write clear technical documentation and benchmark reports for research and engineering audiences.

What We're Looking For
  • 2 to 4 years of experience in research engineering or ML engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.

  • Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.

  • Demonstrated experience designing and running benchmarks or evaluation environments for AI agents or large language models.

  • Experience building infrastructure to reliably run AI models or agents against evaluation tasks at scale.

  • Experience developing metrics or validation studies to assess benchmark difficulty, reliability, and real-world correlation.

  • Ability to collaborate with domain experts and translate complex workflows into evaluation criteria.

  • Strong attention to detail, with a habit of spotting subtle inconsistencies and edge cases.

  • Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces.

  • Excellent written communication skills for technical documentation and cross-timezone collaboration.

  • Published papers or technical writing on AI benchmarking, model evaluation, or failure modes is a strong plus.

  • Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a plus.

Compensation & Benefits

Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.

Location

On-site in Singapore.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmarking Research Engineer - Frontier Agents
AI Benchmarking Research Engineer - Frontier Agents

Clera • Singapore

On-site
SGD 190,000 - 317,000
Visa sponsorship available
Forward Deployed Research Engineer
Forward Deployed Research Engineer

Clera • Singapore

On-site
SGD 190,000 - 317,000
Visa sponsorship
Lead Research Engineer, Data Quality
Lead Research Engineer, Data Quality

Clera • Singapore

On-site
SGD 190,000 - 317,000
Visa sponsorship
Research Engineer, Synthetic Data
Research Engineer, Synthetic Data

Clera • Singapore

On-site
SGD 190,000 - 317,000
Visa sponsorship available
R&D Engineering Manager - AI Evaluation Engine
R&D Engineering Manager - AI Evaluation Engine

Resaro • Singapore

On-site
SGD 150,000 - 210,000
Research Engineer, QC Automation
Research Engineer, QC Automation

Clera • Singapore

On-site
SGD 190,000 - 317,000
Research Engineer Remote/Flexible →
Research Engineer Remote/Flexible →

Neo Research • Singapore

Hybrid
SGD 152,000 - 228,000
Competitive compensation
Flexible work arrangements
(Lead) Research Engineer, Digital Services, IAIC
(Lead) Research Engineer, Digital Services, IAIC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 140,000 - 210,000
(Senior) Research Engineer, Digital Services, IAIC
(Senior) Research Engineer, Digital Services, IAIC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 140,000 - 210,000
AI Product Analyst (Developer Community), AI Verify Foundation
AI Product Analyst (Developer Community), AI Verify Foundation

IMDA • Singapore

On-site
SGD 60,000 - 90,000