An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Nava is building and operating large-scale GPU infrastructure in Bengaluru to power AI workloads. The Compute team handles NVIDIA GPU systems, cluster software, networking, and day-to-day operations to keep thousands of GPUs healthy and responsive for training and inference workloads.
The role requires deep Linux knowledge, NVIDIA stack experience, and leading architecture for compute infrastructure. Hybrid cloud and data center exposure are a plus, with mentoring responsibilities across the
Nava is a neocloud company built for the AI era. We design, deploy, and operate large-scale GPU infrastructure-and deliver inference-as-a-service to teams building next-generation AI products. Our platform runs on NVIDIA GPUs, high-performance networking (like RoCEv2 or InfiniBand), and a fully automated, software-defined operations model. Engineers at Nava work closely with the hardware to keep GPUs fully utilized and models serving efficiently.
Nava is a neocloud company built for the AI era. We design, deploy, and operate large-scale GPU infrastructure-and deliver inference-as-a-service to teams building next-generation AI products. Our platform runs on NVIDIA GPUs, high-performance networking (like RoCEv2 or InfiniBand), and a fully automated, software-defined operations model. Engineers at Nava work closely with the hardware to keep GPUs fully utilized and models serving efficiently.
The Compute team owns Nava's GPU infrastructure end-to-end-from installing and configuring NVIDIA GPU systems, to managing cluster software, networking, and day-to-day operations. Our goal: keep thousands of GPUs healthy, responsive, and running at peak performance for training and inference workloads.
Skills: nvidia,kubernetes,cluster,rdma,infiniband,cuda,architecture,ansible,infrastructure,automation