Our client is one of the most influential technology companies in the world, at the forefront of accelerated computing and AI infrastructure. They are building and operating an enterprise-scale AI Factory - a full-stack environment spanning GPU infrastructure, high-performance networking, distributed storage, and AI platform software that powers some of the most advanced AI workloads on the planet.
We are seeking a Senior Infrastructure Engineer to join this team and own the design, deployment, and operation of large-scale GPU clusters supporting training, fine-tuning, and inference for frontier AI workloads. This role sits at the intersection of HPC, cloud infrastructure, and modern AI platforms - and carries direct responsibility for performance, reliability, and cost efficiency (Perf/TCO) across the entire stack.
If you are the kind of engineer who thinks in clusters, not servers, and who loses sleep over GPU utilization and network saturation before a customer does, we want to talk.
What You'll Be Doing
- Deploy and operate hyperscale GPU clusters (H100/B200-class DGX systems), managing the full cluster lifecycle from provisioning and burn-in through validation, upgrades, and capacity scaling
- Design and operate low-latency, high-bandwidth network fabrics (InfiniBand/RoCE), tuning topology for distributed training workloads including all-reduce and pipeline parallelism
- Build and operate cluster schedulers and orchestrators (Kubernetes and Slurm), implementing multi-tenant workload isolation, quota systems, and utilization optimization across training and inference
- Deploy and optimize high-throughput storage systems including NVMe and parallel file systems (Lustre, GPFS), with a focus on I/O pipeline performance and data locality for GPU-bound workloads
- Drive full-stack Perf/TCO optimization — profiling training and inference workloads to eliminate bottlenecks across GPU utilization, network saturation, and storage throughput
- Build and implement benchmarking frameworks and reproducible performance tests
- Develop deep observability across the hardware and software stack; lead incident response and root-cause analysis across GPUs, nodes, network fabric, and distributed training jobs
What We Need to See
- 8+ years of experience in infrastructure engineering with direct, hands-on experience operating GPU clusters or HPC environments at scale
- Deep expertise in high-performance networking - InfiniBand, RoCE, and RDMA are core to this role, not a nice-to-have
- Production experience with both Kubernetes and Slurm in large-scale AI or HPC environments
- Proficiency with parallel file systems (Lustre, GPFS) and NVMe storage optimization for AI workloads
- Strong performance engineering instincts - you profile, benchmark, and systematically eliminate bottlenecks across the full stack
- Bachelor's degree or higher in Computer Science, Engineering, or a related technical field (or equivalent experience)
- Excellent communication and collaboration skills - this role operates across hardware, software, and research teams
Ways to Stand Out From the Crowd
- Hands-on experience deploying and operating enterprise-grade GPU clusters at production scale
- Experience tuning RDMA software stacks (NCCL, IB verbs, UCX, libfabrics)
- Background in distributed training frameworks and large-scale ML workload patterns (PyTorch, JAX, TensorFlow)
- Experience building observability and telemetry infrastructure across hardware and software layers
- Root cause analysis experience at datacenter scale - across GPU failures, network fabric issues, and distributed training job failures
- Track record of building and owning benchmarking frameworks for reproducible performance testing
Compensation
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base yearly salary range is $150,000 to $350,000.