Senior Research Scientist LLM

techire ai

San Francisco (CA)

On-site

USD 350,000 - 500,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Stock options
Remote work worldwide
Competitive compensation

Job summary

techire ai is a Series A company building the next generation of conversational AI, with an ex-NVIDIA & Meta research leader and a co-creator behind open-source models. The team powers hundreds of millions of conversations monthly, offering work on research that scales to real-world deployments.

They're seeking researchers who have owned meaningful post-training improvements and can tackle data, modelling, evaluation and RL infrastructure challenges to raise model quality and alignment at scale.

Qualifications

  • Experience leading post-training projects delivering tangible capability improvements.
  • Understanding practical challenges behind post-training techniques, not just theory.
  • Experience across data, modelling, evaluation and RL infrastructure, improving model quality.

Responsibilities

  • Design and run post-training experiments using techniques like DPO, GRPO, SFT and rejection sampling.
  • Build and scale reinforcement learning and post-training infrastructure.
  • Develop reward models, evaluation frameworks and assessment rubrics to improve model quality and behavior.
  • Collaborate with vendors and data partners to curate high-quality post-training datasets and feedback pipelines.
  • Own end-to-end projects spanning data collection, modelling, evaluation and infrastructure, driving improvements in reasoning and alignment.

Skills

Post-training projects
RL infrastructure
Evaluation frameworks
Data partnerships
End-to-end project ownership

Job description

The next frontier in LLMs isn't just training models. It's teaching them to reason, use tools, follow intent and improve through post-training.

That's exactly what this team is focused on.

You'll be joining a Series A company building the next generation of conversational AI, working alongside an ex-NVIDIA & Meta research leader and one of the co-creators behind some well known open-source models.

Already experiencing rapid growth with strong commercial traction, their technology powers hundreds of millions of conversations every month, giving you the opportunity to work on research that directly improves models deployed at real-world scale.

They're looking for researchers who've gone beyond implementing methods from papers and have owned meaningful capability improvements themselves.

What you'll do
  • Design and run post-training experiments using techniques such as DPO, GRPO, SFT, rejection sampling and related approaches
  • Build and scale reinforcement learning and post-training infrastructure
  • Develop reward models, evaluation frameworks and assessment rubrics that improve model quality and behaviour
  • Work with vendors and data partners to build high-quality post-training datasets and feedback pipelines
  • Own end-to-end projects spanning data collection, modelling, evaluation and infrastructure, identifying where models break down and driving improvements in reasoning, controllability and alignment
What you'll bring

We're looking for someone who's owned significant post-training projects that have delivered meaningful capability improvements.

You’ll understand the practical challenges behind modern post-training techniques, not just the theory. You've worked across data, modelling, evaluation and RL infrastructure, know how to improve model quality through robust evaluation and preference optimisation, and understand what it takes to make capability improvements repeatable.

Location:

Fully remote worldwide.

Compensation

Depends on experience/location. In SF, comp is up to $500,000 base salary, plus stock.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Senior LLM Researcher - Post-Training & Alignment (Remote)
Senior LLM Researcher - Post-Training & Alignment (Remote)

techire ai • San Francisco (CA)

On-site
USD 350,000 - 500,000
Stock options
Remote work worldwide
Competitive compensation
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

GoTo Meeting • Mountain View (CA)

On-site
USD 150,000 - 230,000
Health, dental, and vision care for you and your family
Top-tier 401(K) plan with company matching
Paid time off and paid holidays
+2
Machine Learning Researcher
Machine Learning Researcher

SOLANA FOUNDATION • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Equity in a high-growth startup
Comprehensive benefits
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Engineer - LLM Post-Training & Agents
Research Engineer - LLM Post-Training & Agents

Kaon (prev. FlowGPT) • San Francisco (CA)

On-site
USD 200,000 - 500,000
Member of ML Technical Staff
Member of ML Technical Staff

Pragmatike • San Francisco (CA)

On-site
USD 200,000 - 350,000
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

NewsBreak • Mountain View (CA)

On-site
USD 130,000 - 160,000
Health, dental, and vision care
401(k) plan with company matching
Paid time off and holidays
Member of Technical Staff - Research, Post-Training
Member of Technical Staff - Research, Post-Training

Modal Labs • New York (NY)

On-site
USD 180,000 - 240,000
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000