Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Together AI seeks an experienced SRE/DevOps engineer to support client-facing Kubernetes GPU clusters in a production environment. The role focuses on high-availability, monitoring, and rapid incident response for AI inference and fine-tuning workloads.
You will work to ensure stability of GPU clusters, collaborate with customers, and guide best practices for deploying and operating AI infrastructure at scale.
Together AI seeks an experienced SRE/DevOps engineer to support client-facing Kubernetes GPU clusters in a production environment. The role focuses on high-availability, monitoring, and rapid incident response for AI inference and fine-tuning workloads.
You will work to ensure stability of GPU clusters, collaborate with customers, and guide best practices for deploying and operating AI infrastructure at scale.