Research Engineer - LLM Post-Training & Agents

Kaon (prev. FlowGPT)

San Francisco (CA)

On-site

USD 200,000 - 500,000

Full time

41 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Kaon, based in the San Francisco Bay Area, is seeking a Research Engineer to advance the models behind our interactive storytelling and character experiences. You will focus on post-training methods to boost model behavior and deploy capable agents with memory, tools, and long-horizon interactions.

In this role, you will build training pipelines (SFT, DPO, GRPO), craft datasets, design reward models, and evaluate improvements through offline tests and live experiments, collaborating with

Qualifications

  • Experience training or post-training large language models, including RL or preference optimization.
  • Ability to turn open-ended research problems into clear hypotheses, experiments, and measurable results.
  • Strong Python and software engineering fundamentals; C/C++ is a plus.

Responsibilities

  • Build and refine LLM post-training pipelines (SFT, DPO, GRPO).
  • Develop training and preference datasets with held-out evaluations.
  • Design reward models for narrative quality, memory, and personalization.
  • Create agent training environments and evaluations for tool use and memory.
  • Collaborate with engineering and product teams to bring improvements into production.
  • Conduct offline evaluations and online A/B tests.

Skills

Python
C/C++
Post-training LLMs
RL/Pref optimization
LLM-based agents
Hypothesis testing
Ownership

Tools

PyTorch
TensorFlow
Linux/UNIX

Job description

Research Engineer – LLM Post-Training & Agents

Kaon (https://www.kaon.io/) builds LLM systems for roleplay, interactive storytelling, and personalized character experiences used by millions of people. We’re looking for a research engineer to improve the models behind these experiences and turn those improvements into capable agents.

Post-training is the core of this role: developing training methods, data, rewards, and evaluations that improve model behavior. You’ll apply that work to agents with strong character consistency, long-term memory, tool use, and coherent multi-turn interactions.

What you’ll work on
  • Build and iterate on LLM post-training pipelines, including supervised fine-tuning, preference optimization, and reinforcement learning (SFT, DPO, GRPO, and related methods).
  • Develop training and preference datasets from user interactions, with careful data quality controls and held-out evaluations.
  • Design reward models and feedback signals for narrative quality, instruction following, character consistency, personalization, and memory accuracy.
  • Build agent training environments and evaluations for tool use, memory, and long-horizon interactions.
  • Explore memory architectures, context management, and memory consolidation that improve an agent’s behavior over time.
  • Run controlled experiments, analyze failures, and validate improvements through offline evaluations and online A/B tests.
  • Work with the engineering and product teams to bring research improvements into production.
What we’re looking for
  • Strong Python skills and solid software engineering fundamentals; C/C++ experience is a plus.
  • Hands‑on experience training or post‑training large language models, including practical familiarity with reinforcement learning or preference optimization.
  • Experience building or evaluating LLM‑based agents, or a strong interest backed by relevant projects.
  • Ability to turn an open‑ended research problem into a clear hypothesis, an experiment, and a measurable result.
  • Ownership from implementation and debugging through evaluation and deployment.
Nice to have
  • Work on agent memory, personalized generation, continual learning, reward modeling, or long‑context evaluation.
  • Experience with roleplay, virtual characters, interactive storytelling, or conversational AI.
  • Research publications or open‑source contributions in relevant areas.

Location: San Francisco Bay Area, on‑site.

Employment: Full‑time.

Compensation: $200,000–$500,000 total compensation (base + equity), depending on experience and impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Post-Training & Agents Research Engineer
LLM Post-Training & Agents Research Engineer

Kaon (prev. FlowGPT) • San Francisco (CA)

On-site
USD 200,000 - 500,000
Member of Technical Staff [Research]
Member of Technical Staff [Research]

NeoCognition Inc. • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Member of Technical Staff, Post-training
Member of Technical Staff, Post-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Senior Research Scientist LLM
Senior Research Scientist LLM

techire ai • San Francisco (CA)

On-site
USD 350,000 - 500,000
Stock options
Remote work worldwide
Competitive compensation
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

GoTo Meeting • Mountain View (CA)

On-site
USD 150,000 - 230,000
Health, dental, and vision care for you and your family
Top-tier 401(K) plan with company matching
Paid time off and paid holidays
+2
Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Researcher
Machine Learning Researcher

SOLANA FOUNDATION • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff, Post-training San Jose
Member of Technical Staff, Post-training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Research Engineer, Post-Training
Research Engineer, Post-Training

Harvey • San Francisco (CA)

On-site
USD 231,000 - 340,000
Agent Post-Training, Artifacts Research
Agent Post-Training, Artifacts Research

OpenAI • San Francisco (CA)

On-site
USD 380,000 - 500,000