Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Thinking Machines is seeking a researcher to own the boundary between reinforcement learning algorithms and the underlying systems. The role focuses on high-training-compute, long-horizon RL, with a center of gravity around asynchronous RL and integration with inference constraints.
You will co-design the RL recipe and the systems, advance async RL, and run frontier-scale RL end-to-end. Candidates should have strong Python and DL framework experience, and a PhD or equivalent research background
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
Our team scales reinforcement learning for frontier models. Progress in RL is increasingly set by how well it scales: more rollouts, larger models, and training loops that keep large fleets of accelerators doing useful work. We are particularly interested in people working on high-training-compute, long-horizon RL. We believe the biggest gains come from designing the training recipe and the infrastructure together rather than separately, and we are hiring a researcher who wants to own that boundary.
A center of gravity for this role is asynchronous RL. Decoupling generation from training changes both the systems design and the learning problem, and doing it well requires a deep understanding of async RL algorithms, design choices, and trade-offs on both the ML and the systems sides. We expect much of the headroom in RL scaling to come from here.
Because generation dominates the cost of RL at scale, good knowledge of inference systems, low-precision numerics, and quantization is recommended: you should be able to reason quantitatively about rollout throughput and cost (batching, KV cache, MoE serving, speculative decoding) and about how inference constraints shape training design.
This is a research role with full-stack ownership, from the algorithms to the parallelism plan to the health of the run.
Minimum qualifications:
Preferred qualifications — we encourage you to apply if you meet some but not all of these:
As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.