Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Cloudjobs is seeking an infrastructure-focused engineer to design, implement, and operate systems for deploying and training AI models on a massive GPU fleet in New York.
You will own job scheduling, cluster provisioning, and high-performance snapshot delivery, collaborating with software and ML teams to ensure scalable, reliable workflows.
Candidate should have strong programming skills and hands-on experience with hyperscale compute, Azure, and Kubernetes; AI/ML workload knowledge is a bonus.
Design, implement, and operate infrastructure systems for model deployment and training on a large-scale GPU fleet. This includes managing job scheduling, cluster provisioning, and high-performance snapshot delivery.
Candidates should have strong programming skills and experience with hyperscale compute systems, specifically using Azure and Kubernetes. An understanding of AI/ML workloads is considered a bonus.