Get more replies from employers
Send a job-specific resume in minutes.
Pareto in San Francisco builds production RL environments and data platforms for frontier labs. You will own environments end to end, from build to production, and work between research teams and engineering to ensure clear specs and reliable delivery.
This role demands shipping fast, owning the release path, and turning training signals into product improvements. Equity is part of the package, and you’ll be at the core of the platform that underpins models at Anthropic and GDM.
Humanity is in a virtuous cycle: human insight improves AI, and better AI expands what people can do. Sustaining it depends on the one input that can't be automated: expert human judgment.
At Pareto, we build the platform that turns that judgment into the data, evals, and RL environments frontier models learn from. We work with leading frontier labs like Anthropic and GDM, and we give skilled people everywhere a way to shape the future of AI and share in what it creates.
This RL environment and human-data infrastructure is already in production. Our job now is to scale it.
You'll own the RL environments frontier labs train on, end to end. Scope the problem with the requester, build the image and the tools inside it, write the graders that score it, ship it into the customers' platform, and keep it healthy once it's running. You sit between Pareto's engineering team and the researchers at the labs we work with, close enough to both that you can tell when a training goal and a buildable spec have drifted apart.
Nobody will hand you a finished spec. You'll get a research problem, define what gets built, and stay with it after it lands. In your first year, good looks like environments that ship faster than the last one did, because you invested in the build and release path instead of hand-rolling each delivery. What you build becomes training signal. That's the reason the ownership runs all the way through production.
Base salary $245K–$300K, plus equity. Final offer depends on experience and level, and we share level-specific ranges early in the process.
The environments you build become the training signal for models at Anthropic and GDM. Not adjacent to that work. Inside it. Equity is part of the package at every level.