Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC

United States

On-site

USD 150,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Summit Group Solutions, LLC is seeking a Senior Infrastructure Engineer to oversee the design, deployment, and operation of large-scale GPU clusters. This role demands expertise in high-performance computing and AI infrastructure, driving performance and cost efficiency across advanced AI workloads. Responsibilities include managing GPU cluster lifecycles, networking design, and performance optimization. A bachelor's degree and 8+ years of engineering experience are necessary. Compensation ranges from $150,000 to $350,000 annually.

Qualifications

  • 8+ years of experience in infrastructure engineering with a focus on GPU clusters or HPC.
  • Deep knowledge of InfiniBand, RoCE, and RDMA.
  • Production experience with Kubernetes and Slurm in large-scale environments.

Responsibilities

  • Deploy and manage hyperscale GPU clusters and their lifecycle.
  • Design low-latency, high-bandwidth network fabrics.
  • Build and operate cluster schedulers and orchestrators.

Skills

Operating GPU clusters
High-performance networking
Kubernetes
Slurm
Performance engineering
Communication skills

Education

Bachelor's degree in Computer Science or related field

Tools

NVMe storage systems
Lustre
GPFS

Job description

Our client is one of the most influential technology companies in the world, at the forefront of accelerated computing and AI infrastructure. They are building and operating an enterprise-scale AI Factory - a full-stack environment spanning GPU infrastructure, high-performance networking, distributed storage, and AI platform software that powers some of the most advanced AI workloads on the planet.

We are seeking a Senior Infrastructure Engineer to join this team and own the design, deployment, and operation of large-scale GPU clusters supporting training, fine-tuning, and inference for frontier AI workloads. This role sits at the intersection of HPC, cloud infrastructure, and modern AI platforms - and carries direct responsibility for performance, reliability, and cost efficiency (Perf/TCO) across the entire stack.

If you are the kind of engineer who thinks in clusters, not servers, and who loses sleep over GPU utilization and network saturation before a customer does, we want to talk.

What You'll Be Doing
  • Deploy and operate hyperscale GPU clusters (H100/B200-class DGX systems), managing the full cluster lifecycle from provisioning and burn-in through validation, upgrades, and capacity scaling
  • Design and operate low-latency, high-bandwidth network fabrics (InfiniBand/RoCE), tuning topology for distributed training workloads including all-reduce and pipeline parallelism
  • Build and operate cluster schedulers and orchestrators (Kubernetes and Slurm), implementing multi-tenant workload isolation, quota systems, and utilization optimization across training and inference
  • Deploy and optimize high-throughput storage systems including NVMe and parallel file systems (Lustre, GPFS), with a focus on I/O pipeline performance and data locality for GPU-bound workloads
  • Drive full-stack Perf/TCO optimization — profiling training and inference workloads to eliminate bottlenecks across GPU utilization, network saturation, and storage throughput
  • Build and implement benchmarking frameworks and reproducible performance tests
  • Develop deep observability across the hardware and software stack; lead incident response and root-cause analysis across GPUs, nodes, network fabric, and distributed training jobs
What We Need to See
  • 8+ years of experience in infrastructure engineering with direct, hands-on experience operating GPU clusters or HPC environments at scale
  • Deep expertise in high-performance networking - InfiniBand, RoCE, and RDMA are core to this role, not a nice-to-have
  • Production experience with both Kubernetes and Slurm in large-scale AI or HPC environments
  • Proficiency with parallel file systems (Lustre, GPFS) and NVMe storage optimization for AI workloads
  • Strong performance engineering instincts - you profile, benchmark, and systematically eliminate bottlenecks across the full stack
  • Bachelor's degree or higher in Computer Science, Engineering, or a related technical field (or equivalent experience)
  • Excellent communication and collaboration skills - this role operates across hardware, software, and research teams
Ways to Stand Out From the Crowd
  • Hands-on experience deploying and operating enterprise-grade GPU clusters at production scale
  • Experience tuning RDMA software stacks (NCCL, IB verbs, UCX, libfabrics)
  • Background in distributed training frameworks and large-scale ML workload patterns (PyTorch, JAX, TensorFlow)
  • Experience building observability and telemetry infrastructure across hardware and software layers
  • Root cause analysis experience at datacenter scale - across GPU failures, network fabric issues, and distributed training job failures
  • Track record of building and owning benchmarking frameworks for reproducible performance testing
Compensation

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base yearly salary range is $150,000 to $350,000.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior AI Network Engineer - AI Infrastructure
Senior AI Network Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
Senior SRE - AI Infrastructure
Senior SRE - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
IPO Equity
10% comapny bonus
401K 4% match