Stand out for this role — generate a tailored resume and cover letter in about a minute.
Recrew AI is looking for a Cluster DevOps/SRE Engineer based in Bengaluru to design and manage GPU/CPU compute clusters on GCP. The role involves ensuring the reliability and performance of critical infrastructure for AI model training workloads.
Candidates should have 4–10 years of experience in DevOps or SRE roles and a strong background in managing large-scale GKE clusters. This position offers competitive compensation and an innovation-driven culture.
Function: DevOps / Site Reliability Engineering / Infrastructure
Type: Full-time
Industry: Artificial Intelligence, Cloud Infrastructure, Telecommunications
The company is the dedicated AI research and innovation arm of a large-scale Indian telecom and technology conglomerate. It is building foundational AI technologies designed specifically for India's languages and digital economy.
Focus areas include multilingual speech recognition, voice synthesis, real-time conversational AI, and multimodal foundation models. The company collaborates with global AI leaders including OpenAI, Anthropic, Google, and Meta.
Its systems are built to serve hundreds of millions of users across India. It is investing across the full AI value chain — from next-generation data centers to consumer AI platforms and an agentic marketplace for SMEs.
The company is seeking a Cluster DevOps / SRE Engineer to design, manage, and optimize large-scale GPU/CPU compute clusters on GCP that power its AI research and model training workloads. The role owns the reliability, performance, and scalability of infrastructure critical to training and serving frontier AI models, sitting at the heart of the company's compute backbone.