Get more replies from employers
Send a job-specific resume in minutes.
STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads.
You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking, and software layers to push performance and scalability. This role requires deep Linux expertise, hands-on container workloads with Pyxis/Enroot, and the ability to implement repeatable benchmarking and automation.
STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads.
You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking, and software layers to push performance and scalability. This role requires deep Linux expertise, hands-on container workloads with Pyxis/Enroot, and the ability to implement repeatable benchmarking and automation.