Get more replies from employers
Send a job-specific resume in minutes.
Career Techniques in New York seeks an experienced infrastructure engineer to design, deploy, and scale large-scale GPU clusters for AI research. You will work across compute, storage, OS, and automation to support hundreds of petabytes and thousands of nodes.
You will profile GPU workloads, remove bottlenecks, and collaborate with researchers to translate findings into speedups. Expect to own end-to-end infrastructure projects from design through long-term support and vendor engagement.
As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes. The role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention.
Comp: 200-300K + Bonus