Turn this role into an interview — a resume and cover letter built around what this employer wants.
Nava in Bengaluru is seeking a capable Systems Engineer to design, deploy, and operate GPU infrastructure for AI workloads. You will manage bare‑metal GPU clusters and ensure fast, reliable training and inference at scale.
You will implement OS provisioning, CUDA drivers, GPU Operator, and RDMA networking; write automation with Python, Ansible, and Terraform; run health checks, benchmarks, and documentation; collaborate with cross‑functional teams to meet technical requirements.
Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the bare metal, on infrastructure built to keep GPUs available and models serving.
The Compute team owns Nava's GPU infrastructure end-to-end: from bare-metal provisioning of NVIDIA GPU systems through cluster software, RDMA integration, and production operation. We keep thousands of GPUs healthy and highly utilized so training and inference workloads run fast and reliably.
NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator; RoCEv2 / InfiniBand RDMA fabrics; Kubernetes (networking, storage, security), Slurm; Ansible, Terraform, Python; monitoring with Prometheus, Grafana, Dynatrace, Datadog, or Zabbix; Linux at scale.