Stand out for this role — generate a tailored resume and cover letter in about a minute.
Thinking Machines Lab in San Francisco, CA is hiring a Site Reliability Engineer to own the reliability and performance of post-training and RL training runs. You will work side by side with researchers, diagnose failures in real time, and build tooling to improve automation and resilience across clusters and pipelines.
The role emphasizes production engineering for model training, with on-call rotations and a focus on high-availability infrastructure and observability.
Thinking Machines Lab in San Francisco, CA is hiring a Site Reliability Engineer to own the reliability and performance of post-training and RL training runs. You will work side by side with researchers, diagnose failures in real time, and build tooling to improve automation and resilience across clusters and pipelines.
The role emphasizes production engineering for model training, with on-call rotations and a focus on high-availability infrastructure and observability.