Research Scientist – RL Post-Training for Agents

Rnb Consultancy

San Francisco (CA)

On-site

USD 250,000 - 500,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Visa sponsorship

Job summary

Rnb Consultancy in San Francisco is hiring a brand-new AI lab focused on autonomous agents that pursue long-horizon goals. You will own ambitious bets from the first hypothesis and dataset through deployment, shaping the research agenda from day one.

The team values exceptional ML research and engineering talent with a track record of leading significant projects, and offers a chance to ship real-world results in a frontier research environment.

Qualifications

  • Exceptional ML research and engineering skills.
  • Depth in RL, LLM post-training, reasoning, agents, or long-horizon systems.
  • Ability to build complete systems from hypothesis to production.
  • Led a significant model, agent system, benchmark, paper, or open-source project.

Responsibilities

  • Use RL and other methods to post-train LLM-based and multimodal agents.
  • Develop long-horizon capabilities: planning, memory, error recovery, proactivity, self-improvement.
  • Build environments where agents use computers and tools, plus data pipelines, benchmarks and evals.
  • Enable agents to understand user intent and stay true to it over long runs.
  • Design careful experiments and transfer results to production.

Skills

Reinforcement Learning
LLM post-training
Agents
Long-horizon systems
Memory & context
Eval harnesses
Production systems

Tools

None

Job description

Want your RL research to land in agents that run for days in the real world, not in a paper appendix?

A brand-new AI lab in San Francisco is building autonomous agents that pursue complex goals over very long horizons. The founding team comes from leading frontier AI labs, autonomous-driving and robotics AI, big-tech research and a top quant firm. This is a research seat that builds real systems: you own ambitious bets from the first hypothesis and dataset all the way to a deployed capability.

They’re hiring 4 research scientists.

What you’ll own
  • Using RL, and whatever else works, to post-train LLM-based and multimodal agents
  • Long-horizon capabilities: planning, memory, error recovery, proactivity, self-improvement, and knowing when to escape to a human
  • Building the environments where agents use computers and tools, plus the data pipelines, benchmarks and evals around them
  • Ways for agents to grasp what a user wants and stay true to it over long runs
  • Careful experiment design, and taking what works all the way into production
  • A real say in the research agenda from day one
What you bring
  • Exceptional ML research and engineering skills
  • Depth in at least one of these: RL, LLM post-training, reasoning, agents, computer use, long-horizon systems, memory and context, evals, or how humans and agents work together
  • Good judgment on what to test and when to stop. You build complete systems and move easily between ideas, large-scale experiments and production.
  • You’ve led something significant: a model, an agent system, a benchmark, a paper, an open-source project or a big research bet
  • Roughly 3–6 years in; frontier-lab experience valued, exceptional outliers and senior leads welcome
  • High agency and comfort with uncertain directions
Bonus points
  • Hands-on RL post-training or computer-use work
  • A strong publication, open-source or benchmark record
  • Experience building environments and eval harnesses
  • A spike: olympiad (IOI/IMO), quant, or world-class competitive achievement
What’s in it for you
  • $250k–$500k base + 1–5% equity
  • Your own research bets, end to end, at a lab where results ship
  • Visa sponsorship available; if you’re outside the US, expect to go via an O-1
Good to know
  • Full-time, in person in San Francisco, 9-9-6
  • Process: informal talk with a founder → technical deep-dive on your research → paid 2–3 day work trial in person → offer
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer
Research Engineer

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Member of Technical Staff - Research Engineer, Post-training
Member of Technical Staff - Research Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive cash and equity (>90th pct
Ownership and autonomy
Lunch onsite
+4
Research Engineer – RL Infrastructure & Agent Environments
Research Engineer – RL Infrastructure & Agent Environments

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff, Post-training
Member of Technical Staff, Post-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Research Engineer/Scientist - Human Alignment, Consumer Devices
Research Engineer/Scientist - Human Alignment, Consumer Devices

OpenAI • San Francisco (CA)

On-site
USD 380,000 - 445,000
Researcher, Synthetic RL
Researcher, Synthetic RL

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 160,000
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Machine Learning Research Scientist
Machine Learning Research Scientist

MaC VC • San Mateo (CA)

On-site
USD 100,000 - 150,000
Competitive compensation against major LLM labs
Unlimited PTO
Flexible working arrangements
+1
Research Engineer
Research Engineer

Tessera Labs • New York (NY)

On-site
USD 200,000 - 300,000
Research Scientist, Data
Research Scientist, Data

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Visa sponsorship