An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Andromeda is seeking a Senior Site Reliability Engineer to design, operate, and optimize large-scale GPU infrastructure used for distributed training and inference. You will work directly with customers to push the limits of AI systems and ensure reliable capacity across multiple regions.
You will own GPU cluster architecture, performance, and reliability engineering, with emphasis on network fabrics, observability, automation, and incident leadership in production environments.
Andromeda is seeking a Senior Site Reliability Engineer to design, operate, and optimize large-scale GPU infrastructure used for distributed training and inference. You will work directly with customers to push the limits of AI systems and ensure reliable capacity across multiple regions.
You will own GPU cluster architecture, performance, and reliability engineering, with emphasis on network fabrics, observability, automation, and incident leadership in production environments.