Senior AI GPU Infra Engineer — Performance & Scale

Summit Group Solutions, LLC

United States

On-site

USD 150,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Summit Group Solutions, LLC is seeking a Senior Infrastructure Engineer to oversee the design, deployment, and operation of large-scale GPU clusters. This role demands expertise in high-performance computing and AI infrastructure, driving performance and cost efficiency across advanced AI workloads. Responsibilities include managing GPU cluster lifecycles, networking design, and performance optimization. A bachelor's degree and 8+ years of engineering experience are necessary. Compensation ranges from $150,000 to $350,000 annually.

Qualifications

  • 8+ years of experience in infrastructure engineering with a focus on GPU clusters or HPC.
  • Deep knowledge of InfiniBand, RoCE, and RDMA.
  • Production experience with Kubernetes and Slurm in large-scale environments.

Responsibilities

  • Deploy and manage hyperscale GPU clusters and their lifecycle.
  • Design low-latency, high-bandwidth network fabrics.
  • Build and operate cluster schedulers and orchestrators.

Skills

Operating GPU clusters
High-performance networking
Kubernetes
Slurm
Performance engineering
Communication skills

Education

Bachelor's degree in Computer Science or related field

Tools

NVMe storage systems
Lustre
GPFS

Job description

Summit Group Solutions, LLC is seeking a Senior Infrastructure Engineer to oversee the design, deployment, and operation of large-scale GPU clusters. This role demands expertise in high-performance computing and AI infrastructure, driving performance and cost efficiency across advanced AI workloads. Responsibilities include managing GPU cluster lifecycles, networking design, and performance optimization. A bachelor's degree and 8+ years of engineering experience are necessary. Compensation ranges from $150,000 to $350,000 annually.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Senior AI Infra SRE: GPU Clusters & High-Perf Networking
Senior AI Infra SRE: GPU Clusters & High-Perf Networking

Andromeda • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Senior AI Infrastructure Performance Engineer
Senior AI Infrastructure Performance Engineer

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 210,000
AI/ML HPC Cluster Engineer — Scale GPU-Accelerated Infra
AI/ML HPC Cluster Engineer — Scale GPU-Accelerated Infra

NVIDIA • Colorado

On-site
USD 124,000 - 196,000
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Senior AI Infra Platform Engineer - GPU Scale (Equity)
Senior AI Infra Platform Engineer - GPU Scale (Equity)

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 300,000 - 350,000
Equity
Senior AI Infra Engineer — Scalable GPU Clusters
Senior AI Infra Engineer — Scalable GPU Clusters

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 288,000
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000