Research Engineer, Synthetic Data

Clera

Singapore

On-site

SGD 190,000 - 317,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Visa sponsorship available

Job summary

Clera is seeking a hands-on research engineer to build synthetic data pipelines that turn domain workflows into scalable training tasks for AI agents. You will join a roughly 15-person engineering team of Olympiad medalists and published researchers, working at the frontier of reinforcement learning and AI alignment.

You will design end-to-end pipelines, collaborate with subject-matter experts, develop diverse, realistic outputs, and build tooling to mutate and validate data at scale.

Qualifications

  • 2–4 years in software/ML engineering or AI research.
  • Proficiency in Python and Docker in Linux environments.
  • Experience building end-to-end synthetic data pipelines for AI/ML.
  • Understanding of synthetic data quality criteria and evaluation metrics.

Responsibilities

  • Design and build end-to-end synthetic data pipelines that convert domain workflows into training tasks.
  • Collaborate with subject-matter experts to create synthetic tasks across domains.
  • Develop synthesis methods producing diverse, realistic, and learnable outputs.
  • Build tooling to mutate, validate, and improve synthetic tasks at scale.
  • Analyze model performance on synthetic tasks and identify failure modes.
  • Define metrics to quantify diversity, realism, learnability, and quality.

Skills

Python
Data pipelines
ML infrastructure
End-to-end ownership

Tools

Docker
Linux

Job description

About the Role

This is a hands-on research engineering role focused on building synthetic data pipelines that turn domain-specific workflows into scalable training tasks for AI agents. You will join a roughly 15-person engineering team of Olympiad medalists and published researchers, working at the frontier of reinforcement learning and AI alignment. The work you do here directly expands what AI models are capable of.

What You'll Do
  • Design and build end-to-end synthetic data pipelines that convert domain-specific workflows into structured, realistic, and challenging training tasks.

  • Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.

  • Develop synthetic task generation methods that produce diverse, realistic, and learnable outputs.

  • Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.

  • Analyze model and agent performance on synthetic tasks to understand what the tasks teach and where they break down.

  • Define and implement metrics to quantify synthetic task diversity, realism, learnability, and overall quality.

What We're Looking For
  • 2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with a focus on data pipelines, ML infrastructure, or synthetic data systems.

  • Proficiency in Python and hands-on experience in Linux environments using containerization tools such as Docker.

  • Demonstrated experience applying synthetic data research methods to build end-to-end data generation pipelines for AI or ML applications.

  • Strong understanding of synthetic data quality criteria and evaluation metrics, including diversity, realism, and learnability, along with awareness of their inherent limitations.

  • Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.

  • Proven ability to independently own and deliver technical projects end-to-end with minimal predefined requirements or roadmap.

  • Sharp eye for edge cases, subtle inconsistencies, and quality issues in synthetic or algorithmically generated datasets.

  • Ability to reason from first principles about task design, scoring, and failure modes.

  • Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.

  • Strong communication skills for effective remote collaboration across time zones.

Compensation & Benefits

Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.

Location

On-site in Singapore.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Research Engineer, Data Quality
Lead Research Engineer, Data Quality

Clera • Singapore

On-site
SGD 190,000 - 317,000
Visa sponsorship
Research Engineer, QC Automation
Research Engineer, QC Automation

Clera • Singapore

On-site
SGD 190,000 - 317,000
Principal Software Engineer, AI & Data Platform
Principal Software Engineer, AI & Data Platform

Base Camp • Singapore

Hybrid
SGD 132,000 - 220,000
Forward Deployed Research Engineer
Forward Deployed Research Engineer

Clera • Singapore

On-site
SGD 190,000 - 317,000
Visa sponsorship
Research Engineer, Benchmarks
Research Engineer, Benchmarks

Clera • Singapore

On-site
SGD 190,000 - 317,000
Visa sponsorship available
Senior Data Engineer
Senior Data Engineer

MOL Accessportal Sdn Bhd • Singapore

On-site
SGD 70,000 - 90,000
#EG Data and Artificial Intelligence Scientist
#EG Data and Artificial Intelligence Scientist

NCS Group • Singapore

On-site
SGD 120,000 - 180,000
AI program leadership
Enterprise strategy shaping
Collaborative team
+1
Staff Research Software Engineer, Google Research - Singapore
Staff Research Software Engineer, Google Research - Singapore

GOOGLE ASIA PACIFIC PTE. LTD. • Singapore

On-site
SGD 180,000 - 240,000
(Senior) Research Engineer, Digital Services, IAIC
(Senior) Research Engineer, Digital Services, IAIC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 140,000 - 210,000
Senior AI Engineer
Senior AI Engineer

OCBC (Singapore) • Singapore

On-site
SGD 120,000 - 180,000
Competitive salary
Extensive learning opportunities