Get more replies from employers
Send a job-specific resume in minutes.
STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads.
You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking, and software layers to push performance and scalability. This role requires deep Linux expertise, hands-on container workloads with Pyxis/Enroot, and the ability to implement repeatable benchmarking and automation.
We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.
Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
Build and support production AI infrastructure running hundreds to thousands of GPUs.
Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
Configure and tune distributed AI software stacks including:
Optimize GPU scheduling and resource allocation for both training and inference environments.
Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
Identify performance regressions and troubleshoot distributed training issues at scale.
Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
Work closely with ML engineers to improve training scalability and inference efficiency.
Create automation to deploy, validate, benchmark, and monitor GPU clusters.
Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.