Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Together AI seeks an experienced SRE/DevOps engineer to support client-facing Kubernetes GPU clusters in a production environment. The role focuses on high-availability, monitoring, and rapid incident response for AI inference and fine-tuning workloads.
You will work to ensure stability of GPU clusters, collaborate with customers, and guide best practices for deploying and operating AI infrastructure at scale.
Together AI provides AI infrastructure, and this role focuses on supporting client-facing Kubernetes GPU clusters. Ideal for an SRE or DevOps engineer with deep experience in GPU technologies and HPC environments, who thrives on solving complex technical issues for customers.
This role involves providing technical support for Together AI's GPU clusters, specifically focusing on customer-facing SRE responsibilities. The core work entails resolving complex technical issues that arise within client Kubernetes GPU environments. This includes maintaining cluster stability and ensuring the reliable operation of AI infrastructure used for inference and fine-tuning.
A key aspect of this position is proactive monitoring of the GPU infrastructure's health. When hardware incidents occur, the engineer is responsible for reporting these to clients. The role also encompasses supporting production infrastructure, which includes managing Slurm, performing node repairs, handling migrations, and ensuring the smooth operation of Kubernetes workloads.
Success in this role means ensuring high availability and performance of client GPU clusters, minimizing downtime through effective problem-solving, and providing expert technical guidance. The engineer acts as a critical interface between Together AI's infrastructure and its customers, directly impacting client satisfaction and operational efficiency of their AI/ML workloads.