RL Research Scientist: Reasoning & Autoformalization

Pramaana Labs

Palo Alto (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Pramaana Labs, a Palo Alto AI lab, is seeking an RL Post-Training Researcher to scale reasoning capabilities of foundation models. You will apply RLVR to new domains, use Lean-derived rewards, and push autoformalization and proving through novel RL algorithms and data evals.

You will own the full loop from rollout to evaluation. Ideal candidates will have deep RL research experience at scale, expertise with deterministic reward signals, and a strong systems mindset to build robust evals with

Qualifications

  • Deep, hands-on research experience in reinforcement learning applied to reasoning models at scale.
  • Experience working with RLVR, test-time RL, or exact deterministic reward signals.
  • Strong algorithmic and systems intuition — comfortable writing custom RL loops, managing data pipelines, and building robust evals from scratch.
  • High autonomy: the ability to take a fuzzy problem area and independently drive it to state-of-the-art results without day-to-day direction.

Responsibilities

  • Scale RLVR to new, real-world domains using the Lean kernel as a deterministic reward signal.
  • Design and implement novel RL algorithms and test-time compute optimizations tailored for formal environments and proof search.
  • Push the frontier of autoformalization, training models to map highly technical natural language into strict formal specifications.
  • Curate data and build evals that tightly correlate with verifiable downstream reasoning capabilities.
  • Own the whole loop: rollout sampling, reward design, the update, the eval.

Skills

Reinforcement learning
RLVR
Custom RL loops

Tools

Lean
Coq
Isabelle

Job description

Pramaana Labs, a Palo Alto AI lab, is seeking an RL Post-Training Researcher to scale reasoning capabilities of foundation models. You will apply RLVR to new domains, use Lean-derived rewards, and push autoformalization and proving through novel RL algorithms and data evals.

You will own the full loop from rollout to evaluation. Ideal candidates will have deep RL research experience at scale, expertise with deterministic reward signals, and a strong systems mindset to build robust evals with

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Reinforcement Learning
Research Scientist, Reinforcement Learning

Pramaana Labs • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Research Engineer/Research Scientist, RL/Reasoning
Research Engineer/Research Scientist, RL/Reasoning

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 160,000
Relocation assistance
Hybrid work model
Commitment to accommodation for disabilities
Senior Research Scientist, Long-Horizon AI & RL (Equity)
Senior Research Scientist, Long-Horizon AI & RL (Equity)

techire ai • San Francisco (CA)

On-site
USD 400,000 - 450,000
Research Engineer/Research Scientist, RL/Reasoning
Research Engineer/Research Scientist, RL/Reasoning

Slope • San Francisco (CA)

On-site
USD 310,000 - 460,000
AI Research Intern: Formal Reasoning & Evaluation
AI Research Intern: Formal Reasoning & Evaluation

Pramaana Labs • Palo Alto (CA)

On-site
USD 40,000 - 65,000
Research Engineer
Research Engineer

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
RL Research Scientist - Post-Training on LLMs & Code Models
RL Research Scientist - Post-Training on LLMs & Code Models

AMD • Santa Clara (CA)

On-site
USD 150,000 - 230,000
Benefits at a glance
Research Engineer: RL & Post-Training LLM Systems
Research Engineer: RL & Post-Training LLM Systems

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Research Scientist - Reinforcement Learning
Research Scientist - Reinforcement Learning

Optimized, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 185,000 - 255,000
Senior Research Scientist LLM
Senior Research Scientist LLM

techire ai • San Francisco (CA)

On-site
USD 350,000 - 500,000
Stock options
Remote work worldwide
Competitive compensation