A complete application in a minute — tailored resume and cover letter, ready to send.
Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to drive the reliability of Tinker end-to-end, collaborating with platform engineers and researchers to make every layer of the system robust.
The role emphasizes defining end-to-end reliability, building observability, incident response, and multi-tenant isolation for large-scale distributed training workloads on Kubernetes.
Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to drive the reliability of Tinker end-to-end, collaborating with platform engineers and researchers to make every layer of the system robust.
The role emphasizes defining end-to-end reliability, building observability, incident response, and multi-tenant isolation for large-scale distributed training workloads on Kubernetes.