A complete application in a minute — tailored resume and cover letter, ready to send.
Mechanize, a San Francisco–based start-up building reinforcement learning environments for frontier AI labs, seeks engineers to design, build, and refine RL tasks. You will own the full lifecycle from ideation to grading and iteration, tackling multi-step workflows, large codebases, and real stakeholder interactions.
You will guide coding agents, evaluate outputs, and contribute to shared infrastructure. You should bring deep software engineering experience across domains and strong Python
Mechanize builds reinforcement learning environments that frontier AI labs use to train and evaluate their coding models. Learn more at mechanize.work.
AI models have gotten good at narrow coding tasks but still fail at the complex, judgment-heavy parts of software engineering. We build the environments that expose those failures and help models improve.
You'll design, build, and refine RL tasks, owning the full lifecycle from ideation through grading, failure analysis, and iteration. At this level, we expect you to work on our most complex tasks: environments involving multi-step workflows, realistic stakeholder interactions, large codebases with real conventions and technical debt, or challenging system design problems.
You will use coding agents heavily, and a large part of the job is directing them well, evaluating their output, and knowing when they are failing in subtle ways. You will also contribute to shared infrastructure and tooling, and may take on mentorship responsibilities for newer team members.
Deep software engineering experience across multiple domains, combined with a strong intuition for AI model behavior. You need to anticipate where a model will take shortcuts, distinguish genuine capability gaps from grader issues, and design tasks that target deeper, more subtle failure modes from areas you know well: infrastructure, distributed systems, performance, security, or other specializations.
This is independent, high-ownership work. You own your tasks from start to finish, with regular feedback. Strong performers are recognized and rewarded. Benefits include health, dental, vision, and life insurance. Applying takes less than one minute.
Interview process: how-our-interview-process-works
Learn more about the work: what-working-here-is-like
~20 person team in San Francisco. Backed by Patrick Collison, Nat Friedman, Daniel Gross, Jeff Dean, Dwarkesh Patel, and Sholto Douglas. Featured in the New York Times, the Dwarkesh Podcast and Hard Fork.