Senior AI GPU Infra Engineer — Performance & Scale
Summit Group Solutions, LLC
United States
On-site
USD 150,000 - 350,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Summit Group Solutions, LLC is seeking a Senior Infrastructure Engineer to oversee the design, deployment, and operation of large-scale GPU clusters. This role demands expertise in high-performance computing and AI infrastructure, driving performance and cost efficiency across advanced AI workloads. Responsibilities include managing GPU cluster lifecycles, networking design, and performance optimization. A bachelor's degree and 8+ years of engineering experience are necessary. Compensation ranges from $150,000 to $350,000 annually.
Qualifications
8+ years of experience in infrastructure engineering with a focus on GPU clusters or HPC.
Deep knowledge of InfiniBand, RoCE, and RDMA.
Production experience with Kubernetes and Slurm in large-scale environments.
Responsibilities
Deploy and manage hyperscale GPU clusters and their lifecycle.
Build and operate cluster schedulers and orchestrators.
Skills
Operating GPU clusters
High-performance networking
Kubernetes
Slurm
Performance engineering
Communication skills
Education
Bachelor's degree in Computer Science or related field
Tools
NVMe storage systems
Lustre
GPFS
Job description
Summit Group Solutions, LLC is seeking a Senior Infrastructure Engineer to oversee the design, deployment, and operation of large-scale GPU clusters. This role demands expertise in high-performance computing and AI infrastructure, driving performance and cost efficiency across advanced AI workloads. Responsibilities include managing GPU cluster lifecycles, networking design, and performance optimization. A bachelor's degree and 8+ years of engineering experience are necessary. Compensation ranges from $150,000 to $350,000 annually.