Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Reflection is building a Kubernetes-first, multi-cloud compute platform to enable scalable GPU workloads and fault-tolerant operations. You will help design and maintain tooling for automatic remediation, capacity planning, and performance debugging across large GPU fleets.
The role emphasizes integrating with training teams to co-design fault tolerance and remediation strategies, while aligning with a multi-cloud storage and data replication roadmap.
Reflection is building a Kubernetes-first, multi-cloud compute platform to enable scalable GPU workloads and fault-tolerant operations. You will help design and maintain tooling for automatic remediation, capacity planning, and performance debugging across large GPU fleets.
The role emphasizes integrating with training teams to co-design fault tolerance and remediation strategies, while aligning with a multi-cloud storage and data replication roadmap.