Research Engineer - Post training & RL

techire ai

California (MO)

Hybrid

USD 180,000 - 300,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
401k
Unlimited PTO
Relocation sponsorship

Job summary

Techire AI is seeking a research-focused candidate to push the boundaries of post-training evaluation for LLMs. You’ll bridge academic insights with practical deployments, drafting papers and implementing scalable evaluation frameworks in live systems.

Location: San Francisco, on-site or hybrid. Salary up to $300K base (DOE) plus meaningful equity and comprehensive benefits including 401k, unlimited PTO, relocation and sponsorship.

Qualifications

  • Research experience in post-training, reinforcement learning, or evaluation for LLMs.
  • Strong understanding of transformer models and experimental design.
  • Publication record at leading venues (NeurIPS, ICLR, ICML, ACL, EMNLP).

Responsibilities

  • The work blends deep research with hands-on implementation in live systems.
  • Publish and defend research findings across venues.
  • Develop evaluation frameworks and simulations to measure progress.

Skills

Post-training RL research
Transformer models
Academic publishing

Education

PhD or equivalent research experience in CS, ML, NLP, or RL

Job description

Want to build the simulated worlds that test what frontier models are really capable of?

This is a chance to join a teamadvancing the science of post-training and scalable evaluation — building reinforcement learning environments that push reasoning, planning, and long-horizon behaviour to their limits.

Instead of static benchmarks, you’ll create dynamic simulations that measure real intelligence — not just accuracy. You’ll design new post-training algorithms (RLHF, DPO, GRPO and beyond), develop richer reward models that move past exact-match scoring, and build evaluation frameworks that define how next-generation AI is trained, aligned, and understood.

The work combines deep research with hands‑on implementation — from writing papers to seeing your methods deployed in live systems. It’s ideal for researchers who care about bridging academic insight and practical impact, helping AI progress beyond metrics that no longer tell the whole story.

You’ll bring:
  • Research experience in post-training, reinforcement learning, or evaluation for LLMs.

  • Strong understanding of transformer models and experimental design.

  • Publication record at leading venues (NeurIPS, ICLR, ICML, ACL, EMNLP).

  • PhD or equivalent research experience in CS, ML, NLP, or RL.

Package: Up to $300K base (DOE) + meaningful equity + comprehensive benefits (401k, unlimited PTO, relocation and sponsorship available).
Location: On-site or hybrid San Francisco.

If you want to shape how AI is trained, tested, and trusted — this is the place to do it.
All applicants will receive a response.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Applied AI
Research Engineer, Applied AI

HeyMilo AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Engineer
Research Engineer

Barrington James • San Francisco (CA)

On-site
USD 140,000 - 210,000
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research, Post-Training
Research, Post-Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research Engineer – RL & Post-Training Evaluation (Equity)
Research Engineer – RL & Post-Training Evaluation (Equity)

techire ai • California (MO)

Hybrid
USD 180,000 - 300,000
Equity
401k
Unlimited PTO
+1
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Engineer/Scientist - Human Alignment, Consumer Devices
Research Engineer/Scientist - Human Alignment, Consumer Devices

OpenAI • San Francisco (CA)

Hybrid
USD 380,000 - 445,000
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Research Scientist - Long-Horizon Multi-Agent Systems
Research Scientist - Long-Horizon Multi-Agent Systems

techire ai • San Francisco (CA)

On-site
USD 400,000 - 450,000