An application made for this job — a tailored resume and cover letter that speak straight to the posting.
GMI Cloud is seeking a Site Reliability Engineer to join the Global Infrastructure team in the United States. This hands-on role ensures the stability, efficiency, and reliability of large-scale AI/ML clusters in our data centers.
You will design, deploy, and maintain scalable AI infrastructure, monitor GPU cluster health, and automate deployment and lifecycle management of GPU nodes. Experience with automation tools and Kubernetes is valued.
GMI Cloud is seeking a Site Reliability Engineer to join the Global Infrastructure team in the United States. This hands-on role ensures the stability, efficiency, and reliability of large-scale AI/ML clusters in our data centers.
You will design, deploy, and maintain scalable AI infrastructure, monitor GPU cluster health, and automate deployment and lifecycle management of GPU nodes. Experience with automation tools and Kubernetes is valued.