Stand out for this role — generate a tailored resume and cover letter in about a minute.
GMI Cloud, a fast-growing AI infrastructure startup in Silicon Valley, seeks a dynamic Site Reliability Engineer with expertise in Kubernetes. This hands-on role is essential for managing the stability and efficiency of high-performance AI/ML clusters.
The ideal candidate will have a Bachelor's degree in Computer Science and over 3 years of experience in data center operations. Responsibilities include designing scalable solutions, monitoring system performance, and automating infrastructure resource management.
We are a fast-growing AI infrastructure startup based in Silicon Valley, working on cutting-edge technologies that power the future of artificial intelligence.We power developers, startups, and enterprises with scalable GPU cloud and inference solutions, helping AI builders turn ideas into reality. As we expand globally, we are looking for a dynamic and hands-on Site Reliability Engineer
We are seeking a skilled Site Reliability Engineer in the area of Kubernetes to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.