An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Nava designs, deploys, and operates large-scale GPU infrastructure and delivers inference-as-a-service for AI products. The Compute team owns GPU systems, software, networking, and day-to-day operations to keep thousands of GPUs healthy and running at peak performance.
We seek an experienced infrastructure engineer to architect, implement, and automate tasks across GPU clusters, leveraging RDMA, NVLink, and the NVIDIA stack for reliable, scalable AI workloads.
Nava is a neocloud company built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building next-generation AI products. Our platform runs on NVIDIA GPUs, high-performance networking (like RoCEv2 or InfiniBand), and a fully automated, software-defined operations model. Engineers at Nava work closely with the hardware to keep GPUs fully utilized and models serving efficiently.
The Compute team owns Nava's GPU infrastructure end-to-end from installing and configuring NVIDIA GPU systems, to managing cluster software, networking, and day-to-day operations. Our goal: keep thousands of GPUs healthy, responsive, and running at peak performance for training and inference workloads.