Get more replies from employers
Send a job-specific resume in minutes.
Andromeda Cluster is seeking a Senior Site Reliability Engineer to design, operate and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems.
You will own GPU cluster architecture, optimize performance, ensure reliability, and build automation and observability tooling. The role requires hands-on experience with GPU hardware, Linux internals, Kubernetes and high-speed interconnects.
Andromeda Cluster is seeking a Senior Site Reliability Engineer to design, operate and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems.
You will own GPU cluster architecture, optimize performance, ensure reliability, and build automation and observability tooling. The role requires hands-on experience with GPU hardware, Linux internals, Kubernetes and high-speed interconnects.