A complete application in a minute — tailored resume and cover letter, ready to send.
Calance seeks a Senior AI Infrastructure Engineer to own the scalable GPU training pipeline, from bring-up to operation. You will design fault-tolerant, self-healing GPU clusters, tune high-speed interconnects, and automate deployment with IaC.
You will enhance NCCL, Run:AI, and Ray scheduling, ensuring high availability for ML workloads across teams. You will lead hardware bring-up, firmware management, and fabric optimization while maintaining fault isolation and observability in a hybrid
Calance seeks a Senior AI Infrastructure Engineer to own the scalable GPU training pipeline, from bring-up to operation. You will design fault-tolerant, self-healing GPU clusters, tune high-speed interconnects, and automate deployment with IaC.
You will enhance NCCL, Run:AI, and Ray scheduling, ensuring high availability for ML workloads across teams. You will lead hardware bring-up, firmware management, and fabric optimization while maintaining fault isolation and observability in a hybrid