An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Pramaana Labs, a Palo Alto AI lab, is seeking an RL Post-Training Researcher to scale reasoning capabilities of foundation models. You will apply RLVR to new domains, use Lean-derived rewards, and push autoformalization and proving through novel RL algorithms and data evals.
You will own the full loop from rollout to evaluation. Ideal candidates will have deep RL research experience at scale, expertise with deterministic reward signals, and a strong systems mindset to build robust evals with
Pramaana Labs, a Palo Alto AI lab, is seeking an RL Post-Training Researcher to scale reasoning capabilities of foundation models. You will apply RLVR to new domains, use Lean-derived rewards, and push autoformalization and proving through novel RL algorithms and data evals.
You will own the full loop from rollout to evaluation. Ideal candidates will have deep RL research experience at scale, expertise with deterministic reward signals, and a strong systems mindset to build robust evals with