A complete application in a minute — tailored resume and cover letter, ready to send.
Sigma Nova is seeking a Research Scientist (Reinforcement Learning) to bring RL into the core of how foundation models are adapted for industrial use. You’ll design learning signals from expert feedback and build robust evaluation protocols.
You’ll work with engineers to deploy RL methods on production-ready models, address safety and uncertainty, and publish open research artifacts as appropriate.
We’re hiring a Research Scientist (Reinforcement Learning) to help bring RL into the core of how foundation models are adapted and improved for industrial use.
In industry, models don’t live in isolation: domain experts validate, correct, and act on model outputs. We want RL to leverage that expert feedback not only as a post‑training patch, but increasingly earlier in the pipeline, shaping objectives, training signals, and adaptation strategies.
Turn expert validation/correction into a reliable learning signal.
Design feedback interfaces/signals that are practical in real operational settings.
Develop RL methods that sit on top of (or integrate with) foundation models used in production.
Explore ways for RL to intervene earlier in the chain (not just after deployment).
Build evaluation protocols aligned with real constraints: robustness, uncertainty reduction, safety, auditability, and cost of error.
Work closely with scientists/engineers to ship demonstrators that connect benchmarks to field outcomes.
PhD (preferred) or equivalent research experience in Reinforcement Learning / Machine Learning.Strong foundations in RL (e.g., policy optimization, off‑policy learning, offline RL, exploration, credit assignment).Ability to design rigorous experiments, debug failure modes, and iterate fast with scientific discipline.Strong programming skills (Python; deep learning stack such as PyTorch).Nice to haveExperience with real‑world RL constraints (noisy/limited feedback, safety requirements, deployment considerations).Comfort with complex data modalities (time series, scientific/industrial signals, multimodal setups).Publications or open research artifacts in RL / sequential decision‑making.
Work on RL problems that matter in the real world: expert feedback loops, uncertainty reduction, and mission‑critical constraints.A research culture that values clarity, rigor, and humility and that connects fundamental ideas to deployable systems.High ownership in a small team: you’ll shape direction, not just execute tasks.