Stand out for this role — generate a tailored resume and cover letter in about a minute.
Thinking Machines is seeking a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You will work alongside the platform and research teams to make every layer of the system more robust and resilient.
You will define and own end-to-end reliability, develop SLOs for distributed training systems, and design monitoring and incident response across the full training path, ensuring secure, multi-tenant operation at scale.
Thinking Machines is seeking a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You will work alongside the platform and research teams to make every layer of the system more robust and resilient.
You will define and own end-to-end reliability, develop SLOs for distributed training systems, and design monitoring and incident response across the full training path, ensuring secure, multi-tenant operation at scale.