Get more replies from employers
Send a job-specific resume in minutes.
Our client, a well-funded AI company, designs and runs large-scale compute infrastructure powering frontier model training and inference. They seek an infrastructure engineer to design, operate, and improve their GPU cluster platform, a deeply technical role at the junction of distributed systems and ML platform engineering.
You will own the compute platform, optimize GPU resource sharing, tackle bottlenecks across compute, storage and networking, and build automation and observability.
Our client is a well-funded AI company operating at significant scale, building and running the large-scale compute infrastructure that powers frontier model training and inference. They are looking for a strong infrastructure engineer to design, operate, and continuously improve their GPU cluster platform. This is a high-impact, deeply technical role sitting at the intersection of distributed systems, ML platform engineering, and large-scale operations — ideal for engineers who want to work close to the metal on some of the most demanding infrastructure challenges in the industry.