Founding AI Research Engineer - AI Evaluation, Benchmarking

LH2 AI Labs

Bengaluru

On-site

INR 1,800,000 - 3,800,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LH2 AI Labs is building post-training infrastructure for frontier AI models. We seek an AI Research Engineer to design evaluation pipelines, create coding benchmarks, and build reproducible data-tasks from real repositories.

You will work at the intersection of software engineering and AI evaluation, crafting test harnesses, verifiers, and datasets for RLHF, RLVR, and other post-training workflows.

Qualifications

  • 3-6+ years of professional experience in software engineering, AI research engineering, machine-learning engineering, or developer tooling.
  • ,

Responsibilities

  • Design realistic coding tasks covering debugging, feature development, refactoring, code review, repository navigation, and test-based problem solving.
  • Source and structure tasks from repositories, issues, pull requests, commits, diffs, test suites, and customer-provided codebases.
  • Build reproducible evaluation environments using repositories, Docker containers, dependencies, test harnesses, sandboxes, and code-execution workflows.
  • Define task instructions, acceptance criteria, reference solutions, grading rubrics, tests, hidden tests, and programmatic verifiers.
  • Run and adapt coding benchmarks such as SWE-bench, SWE-Bench Pro, DeepSWE, Terminal-Bench, or similar internal benchmarks.
  • Evaluate outputs and trajectories generated by LLMs, humans, and coding agents, and determine whether tasks have been correctly solved.
  • Identify ambiguous, broken, trivial, contaminated, non-reproducible, or incorrectly verified tasks before they reach customers or training pipelines.
  • Analyse model and agent failure modes across repository understanding, planning, implementation, tool usage, debugging, and test execution.
  • Build datasets for supervised fine-tuning, preference optimisation, RLHF, RLVR, and other LLM post-training workflows.
  • Convert successful and unsuccessful agent trajectories into demonstrations, corrections, preference pairs, critiques, and outcome-labelled training examples.
  • Build and improve pipelines for collecting, validating, deduplicating, analysing, and packaging evaluation and post-training datasets.
  • Work with customers and internal teams to convert high-level evaluation or training goals into clear benchmark workflows, technical environments, and quality standards.

Skills

Software engineering
AI research
Developer tooling
Git workflows
CI pipelines
Docker environments
Benchmarking
Evaluation design
Code reading
Test harnesses
Verifiers
Documentation
Communication

Tools

SWE-bench
SWE-Bench Pro
DeepSWE
Terminal-Bench

Job description

Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.

We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.

For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on- there is limited new signal left there. That is where we come in.

Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today.

The Role

We are looking for an AI Research Engineer - AI Evaluation, Benchmarking & LLM Post-Training to help us build evaluation and training-data systems for frontier AI models and coding agents.

In this role, you will work at the intersection of software engineering, AI evaluation, benchmark design, and post-training. You will create realistic coding tasks from real repositories, build reproducible evaluation environments, design tests and verifiers, analyse model performance, and improve the quality of data used to train and evaluate AI systems.

This is a strong fit for a research-oriented software engineer with practical experience in coding benchmarks such as SWE-bench, SWE-Bench Pro, DeepSWE, Terminal-Bench, or equivalent repository-based evaluation frameworks.

What You’ll Do

  • Design realistic coding tasks covering debugging, feature development, refactoring, code review, repository navigation, and test-based problem solving.
  • Source and structure tasks from repositories, issues, pull requests, commits, diffs, test suites, and customer-provided codebases.
  • Build reproducible evaluation environments using repositories, Docker containers, dependencies, test harnesses, sandboxes, and code-execution workflows.
  • Define task instructions, acceptance criteria, reference solutions, grading rubrics, tests, hidden tests, and programmatic verifiers.
  • Run and adapt coding benchmarks such as SWE-bench, SWE-Bench Pro, DeepSWE, Terminal-Bench, or similar internal benchmarks.
  • Evaluate outputs and trajectories generated by LLMs, humans, and coding agents, and determine whether tasks have been correctly solved.
  • Identify ambiguous, broken, trivial, contaminated, non-reproducible, or incorrectly verified tasks before they reach customers or training pipelines.
  • Analyse model and agent failure modes across repository understanding, planning, implementation, tool usage, debugging, and test execution.
  • Build datasets for supervised fine-tuning, preference optimisation, RLHF, RLVR, and other LLM post-training workflows.
  • Convert successful and unsuccessful agent trajectories into demonstrations, corrections, preference pairs, critiques, and outcome-labelled training examples.
  • Build and improve pipelines for collecting, validating, deduplicating, analysing, and packaging evaluation and post-training datasets.
  • Work with customers and internal teams to convert high-level evaluation or training goals into clear benchmark workflows, technical environments, and quality standards.

What We’re Looking For

  • 3-6+ years of professional experience in software engineering, AI research engineering, machine-learning engineering, or developer tooling.
  • Strong ability to read, understand, modify, and debug unfamiliar production codebases.
  • Experience with Git, GitHub workflows, pull requests, issues, commits, diffs, branches, and code review.
  • Experience working with unit tests, integration tests, CI workflows, failing builds, dependency issues, and debugging logs.
  • Ability to create and debug reproducible environments using Docker, Linux, shell scripts, package managers, and cloud or local execution setups.
  • Hands-on experience building, running, adapting, or analysing coding-model or coding-agent benchmarks.
  • Practical knowledge of benchmarks such as SWE-bench, SWE-bench Verified, SWE-Bench Pro, DeepSWE, Terminal-Bench, or equivalent internal frameworks.
  • Experience creating evaluation tasks from repositories, issues, pull requests, commits, or real engineering requirements.
  • Experience designing test harnesses, hidden tests, reference solutions, grading rubrics, golden datasets, or programmatic verifiers.
  • Understanding of benchmark contamination, solution leakage, test overfitting, flaky environments, and reproducibility challenges.
  • Ability to distinguish between model failures, agent failures, task-design problems, verifier problems, and environment failures.
  • Strong engineering judgment around correctness, reproducibility, code quality, task difficulty, and evaluation reliability.
  • Strong written and verbal communication skills for documenting task design, methodology, results, and failure analysis.

Nice to Have

  • Experience building benchmarks or evaluation datasets for private repositories.
  • Experience with coding agents or developer tools such as SWE-agent, OpenHands, Claude Code, Cursor, Codex-style agents, or Devin-style systems.
  • Experience building datasets for RLHF, DPO, RLVR, verifier-based reinforcement learning, or other post-training workflows.
  • Familiarity with additional benchmarks such as LiveCodeBench, HumanEval+, MBPP+, RepoBench, or Multi-SWE-bench.
  • Experience with secure code execution, sandboxing, automated grading, test generation, or distributed evaluation infrastructure.
  • Prior experience working with AI labs, evaluation teams, data companies, or applied AI organisations.

Success in This Role Looks Like

  • Coding tasks are realistic, reproducible, well-scoped, challenging, and objectively verifiable.
  • Public and private repositories can be converted into reliable benchmark environments.
  • Evaluation results accurately reflect model capabilities rather than problems in the task, verifier, or infrastructure.
  • Broken, ambiguous, contaminated, or low-quality tasks are identified early.
  • Model failures are converted into useful benchmark improvements and high-quality post-training examples.
  • Customers receive clean datasets, clear methodologies, meaningful analysis, and trustworthy evaluation results.

Why This Role Matters

Training and evaluating AI coding systems requires more than collecting code or comparing generated patches with existing Git diffs.

Reliable benchmarks require realistic engineering tasks, reproducible repositories, strong tests, programmatic verifiers, secure execution environments, and careful human judgment.

This role will help LH2 AI Labs build the evaluation and post-training infrastructure required to improve the next generation of frontier AI models and autonomous software-engineering agents.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding AI Research Lead
Founding AI Research Lead

LH2 AI Labs • Bengaluru

On-site
INR 400,000 - 700,000
Founding AI Researcher Lead
Founding AI Researcher Lead

LH2 AI Labs • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Senior AI Evaluation & Reliability Engineer
Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt. Ltd. • Ahmedabad District

On-site
INR 3,000,000 - 6,000,000
Great Place To Work certified
LLM / Agentic Evaluation Rig Engineer
LLM / Agentic Evaluation Rig Engineer

PHIZENIX • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
RL Environment Researcher
RL Environment Researcher

Provue • Mumbai

On-site
INR 1,500,000 - 2,800,000
ML Researcher – Benchmarks & Evaluation
ML Researcher – Benchmarks & Evaluation

Deccan AI Automation Private Limited • Hyderabad

On-site
INR 2,000,000 - 3,600,000
Machine Learning Engineer, Evaluation
Machine Learning Engineer, Evaluation

HackerRank • Bengaluru

On-site
INR 7,554,000 - 9,443,000
Research Engineer — Agent Architectures (Coding & Autonomous Systems)
Research Engineer — Agent Architectures (Coding & Autonomous Systems)

Lexsi Labs • Mumbai

Hybrid
INR 1,200,000 - 2,400,000
Machine Learning Engineer
Machine Learning Engineer

Pilotcrew AI • India

Remote
INR 2,500,000 - 5,000,000
Work on cutting-edge AI projects
High technical ownership
Opportunity to shape AI benchmarks
Senior Software Engineer
Senior Software Engineer

Whereuelevate • Gurgaon

On-site
INR 1,200,000 - 1,800,000
Work on cutting-edge LLM systems
Be part of a high-growth, high-impact team