Stand out for this role — generate a tailored resume and cover letter in about a minute.
Thinking Machines is hiring a Site Reliability Engineer in San Francisco to own the health of post-training and RL pipelines, partnering with researchers to unblock training and harden infrastructure. Youll implement recovery, observability, and tooling to keep runs fast and reliable.
Youll operate in a high-scale ML environment, collaborate across teams, and participate in on-call rotations while driving permanent fixes and reducing toil for researchers.
Thinking Machines is hiring a Site Reliability Engineer in San Francisco to own the health of post-training and RL pipelines, partnering with researchers to unblock training and harden infrastructure. Youll implement recovery, observability, and tooling to keep runs fast and reliable.
Youll operate in a high-scale ML environment, collaborate across teams, and participate in on-call rotations while driving permanent fixes and reducing toil for researchers.