A complete application in a minute — tailored resume and cover letter, ready to send.
Evergrid seeks an experienced HPC GPU operations engineer to own the health, reliability, and performance of AMD GPU clusters. You will lead Linux systems, kernel debugging, and ROCm-based ML stack across production-scale AI workloads.
You will be on-call for outages, optimize GPU topology, and work with data centers and vendors. A strong background in HPC, Python, Bash, and ROCm is required.
Evergrid seeks an experienced HPC GPU operations engineer to own the health, reliability, and performance of AMD GPU clusters. You will lead Linux systems, kernel debugging, and ROCm-based ML stack across production-scale AI workloads.
You will be on-call for outages, optimize GPU topology, and work with data centers and vendors. A strong background in HPC, Python, Bash, and ROCm is required.