RL Systems Engineer for Frontier AI Tinker

Thinking Machines Lab Inc.

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. is hiring to build our training systems for Tinker, including RL systems, numerics, kernels, and beyond. You will work with researchers and external partners to push frontier models and enable post-training capabilities across the stack.

You’ll co-design RL algorithms, diagnose training runs, and optimize pipelines to deliver frontier-level results, making our platform and models perform at peak capability for researchers and developers alike.

Qualifications

  • Bachelor’s degree or equivalent in CS/ML/Physics/Math with strong grounding.
  • Proficiency in Python and DL frameworks; comfort debugging distributed training.
  • Clarity in communication and ability to explain complex technical concepts.
  • Strong interest in working on Tinker and increasing usefulness and adoption.
  • Experience with RL training stability techniques preferred.
  • Familiarity with low-precision training/inference and LLM serving stacks.

Responsibilities

  • Develop frontier customization techniques and build the post-training engine.
  • Engage with researchers and external partners pushing Tinker.
  • Co-design RL algorithms and training systems across the stack.
  • Debug RL runs, optimize post-training pipelines, and improve performance.

Skills

Python
DL frameworks
Distributed training
Communication
Interest in Tinker

Education

Bachelor's degree or equivalent

Tools

PyTorch
TensorFlow
JAX
SGLang
vLLM

Job description

Thinking Machines Lab Inc. is hiring to build our training systems for Tinker, including RL systems, numerics, kernels, and beyond. You will work with researchers and external partners to push frontier models and enable post-training capabilities across the stack.

You’ll co-design RL algorithms, diagnose training runs, and optimize pipelines to deliver frontier-level results, making our platform and models perform at peak capability for researchers and developers alike.

Get your free, confidential resume review.

or drag and drop your file here.