Founding Member of Technical Staff (RL Environments and Evaluations)

LH2 AI Labs

Bengaluru

On-site

INR 1,800,000 - 3,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

LH2 AI Labs seeks a results-driven expert to shape post-training data tasks, verifiers, and reward signals for frontier AI models. You will design environments, evals, and reasoning traces labs will pay for, ensuring robust, hard-to-game datasets across coding, ops, medical, and more.

You will lead post-training experiments, prove genuine learning signals, and stay ahead of evolving frontier-lab directions to drive demand. A strong research-minded, small-team contributor is essential.

Qualifications

  • Hands-on post-training and evaluation experience (SFT, RLHF/RLVR, reward modeling).
  • Experience building or training with these methods (trained models, evals, or RL environments/verifiers).
  • Strong instinct to craft useful, hack-resistant tasks and verifiers.
  • Ability to translate raw data into tasks, graders and measurable value.

Responsibilities

  • Define vertical-specific task quality and the bar for verifiers.
  • Design methods to turn raw data into tasks, evals and traces.
  • Own verifier and reward-function quality, guarding against gaming.
  • Run post-training experiments to prove learning signal and publish benchmarks.
  • Stay ahead of frontier labs' training directions to create demand.
  • Partner with founders on capabilities and verticals to pursue.

Skills

SFT RLHF RLVR
Reward modeling
RL environments
Benchmarks evaluation
Task design & verifiers
Data-to-task translation
Research mindset

Education

Graduate from Tier 1 engineering institution (IIT/BITS/NIT/IIIT or equivalent)

Job description

Mandatory requirement - Graduate from a Tier 1 engineering institution such as IIT, BITS, NIT, IIIT, or equivalent

About LH2 AI Labs

Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.

We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.

For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on—there is limited new signal left there. That is where we come in.

Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today.

About the role:

Own what makes our data valuable to a frontier lab—the design of tasks, the robustness of verifiers and reward signals, and the judgment of what actually moves a model's capability.

What you'll do
  • Define, per vertical, what a good task, a robust verifier, and a valuable dataset are—and set the quality bar everything ships against.
  • Design the methodology for turning raw substrate (codebases, ops data, medical data) into tasks, RL environments, evals, and reasoning traces that labs will pay for.
  • Own verifier and reward-function quality—including hardening against reward hacking, which is where task value is won or lost.
  • Run real post-training/eval experiments on our own data to prove our environments and tasks produce a genuine learning signal and to generate credibility (published benchmarks, measured uplift).
  • Stay ahead of what Frontier Labs are training toward (long-horizon agentic, multi-file coding, domain reasoning) so we build tomorrow's demand, not yesterday's.
  • Be the founder's and product's thought partner on which capabilities and verticals to pursue.
Must-have skills
  • Deep, hands-on understanding of post-training and evaluation—SFT, RLHF/RLVR, reward modeling, RL environments, and benchmarks—not just familiarity.
  • Has actually built or trained with these methods (trained models, designed evals, or built RL environments/verifiers), not only read about them.
  • Strong instinct for what separates a useful, hack-resistant task/verifier from a gameable or trivial one.
  • Ability to translate between raw data and "Here's the task, here's the grader, here's the measured value."
  • Comfort setting and holding a research quality bar in a small team.
Nice-to-have skills
  • Prior frontier-lab or strong applied-research-laboratory experience.
  • Published benchmarks, evals, or open datasets.
  • Domain depth in one of our verticals (coding, company ops, medical).
  • Familiarity with contamination, difficulty calibration, and eval integrity.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Software Engineer - Data Products
Founding Software Engineer - Data Products

LH2 AI Labs • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior AI Evaluation & RLHF Specialist
Senior AI Evaluation & RLHF Specialist

Innodata India • Dadri

On-site
INR 1,800,000 - 2,800,000
AI Technical Lead
AI Technical Lead

Benchmarkit • Pune District

On-site
INR 4,000,000 - 7,000,000
AI Trainer
AI Trainer

DevFixr • Pune District

On-site
INR 1,653,000 - 3,306,000
Strategic Projects Lead
Strategic Projects Lead

Bespoke Labs • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Member of Technical Staff - Post-Training Engineer
Member of Technical Staff - Post-Training Engineer

Landeed | YC • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Newpage - Full stack AI engineer - Python/React.js
Newpage - Full stack AI engineer - Python/React.js

Newpage Solutions • Maharashtra

On-site
INR 2,400,000 - 4,200,000
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)

Solutions By Text • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Senior AI/ML Engineer/ Developer
Senior AI/ML Engineer/ Developer

RADcube • Hyderabad

On-site
INR 2,500,000 - 3,500,000
AI Engineer
AI Engineer

Benchmarkit • Pune District

On-site
INR 600,000 - 1,200,000