Senior AI GPU Cluster Architect

STN Inc

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads.

You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking, and software layers to push performance and scalability. This role requires deep Linux expertise, hands-on container workloads with Pyxis/Enroot, and the ability to implement repeatable benchmarking and automation.

Qualifications

  • 7+ years designing or operating large-scale Linux infrastructure.
  • 5+ years supporting production GPU clusters for AI or HPC workloads.
  • Demonstrated experience building multi-node GPU training environments from the ground up.
  • Deep expertise with distributed PyTorch training.
  • Extensive experience troubleshooting and optimizing NCCL communications.
  • Strong understanding of distributed AI communication patterns: AllReduce, ReduceScatter, AllGather, Broadcast, point-to-point.
  • Experience benchmarking distributed training using nccl-tests, NVIDIA DCGM, Nsight Systems, MLPerf (preferred).
  • Strong understanding of GPU memory management: KV Cache, activation checkpointing, tensor/pipeline/data parallelism.
  • Experience tuning CUDA, NCCL, UCX, and MPI for max distributed performance.
  • Expert-level Linux systems administration skills.
  • Experience with Slurm workload manager.
  • Experience using Pyxis and Enroot for containerized GPU workloads.
  • Strong scripting skills in Python and Bash.

Responsibilities

  • Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
  • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
  • Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
  • Build and support production AI infrastructure running hundreds to thousands of GPUs.
  • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
  • Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
  • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
  • Configure and tune distributed AI software stacks including: PyTorch, NCCL, CUDA, UCX, MPI, Slurm, Pyxis/Enroot.
  • Optimize GPU scheduling and resource allocation for both training and inference environments.
  • Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
  • Identify performance regressions and troubleshoot distributed training issues at scale.
  • Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
  • Work closely with ML engineers to improve training scalability and inference efficiency.
  • Create automation to deploy, validate, benchmark, and monitor GPU clusters.
  • Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.

Skills

GPU clusters
Distributed PyTorch
NCCL
CUDA
UCX
MPI
Slurm
Pyxis/Enroot
Python
Bash
Linux systems administration

Tools

InfiniBand
RoCE v2
RDMA
GPUDirect RDMA
GPUDirect Storage
NVIDIA DCGM
Nsight Systems
nccl-tests
MLPerf

Job description

STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads.

You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking, and software layers to push performance and scalability. This role requires deep Linux expertise, hands-on container workloads with Pyxis/Enroot, and the ability to implement repeatable benchmarking and automation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
GPU Infra Solutions Architect for Large-Scale AI Clusters
GPU Infra Solutions Architect for Large-Scale AI Clusters

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Senior AI Infrastructure Engineer - GPU Compute
Senior AI Infrastructure Engineer - GPU Compute

Unchain Data • United States

On-site
USD 120,000 - 160,000