Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Prime Intellect AI in San Francisco seeks an experienced engineer to design GPU cluster architectures, deploy strategies for LLM training and HPC workloads, and present architectural recommendations to clients.

You will deploy orchestration with SLURM and Kubernetes, optimize interconnects, and provide 24/7 on‑call support for critical deployments, collaborating with top engineering teams to push frontier AI infrastructure.

Qualifications

  • 3+ years hands-on experience with GPU clusters and HPC environments.
  • Deep expertise with SLURM and Kubernetes in production GPU settings.
  • Proven experience with InfiniBand configuration and troubleshooting.
  • Strong understanding of NVIDIA GPU architecture, CUDA ecosystem, and driver stack.
  • Experience with infrastructure automation tools (Ansible, Terraform).
  • Proficiency in Python, Bash, and systems programming.
  • Track record of customer-facing technical leadership.

Responsibilities

  • Partner with clients to understand workloads and design GPU cluster architectures.
  • Create proposals and capacity plans for clusters from 100 to 10,000+ GPUs.
  • Develop deployment strategies for LLM training, inference, and HPC workloads.
  • Present architectural recommendations to technical and executive stakeholders.
  • Deploy and configure SLURM and Kubernetes for distributed workloads.
  • Implement high-performance networking and optimize GPU utilization.
  • Provide 24/7 on-call support for critical deployments.
  • Create runbooks and documentation for operations teams.
  • Diagnose and resolve complex hardware, drivers, networking, and software problems.

Skills

GPU clusters
HPC environments
SLURM
Kubernetes
NVIDIA CUDA
Python
Bash
Customer-facing leadership

Tools

Ansible
Terraform
Docker
Containerd

Job description

Prime Intellect AI in San Francisco seeks an experienced engineer to design GPU cluster architectures, deploy strategies for LLM training and HPC workloads, and present architectural recommendations to clients.

You will deploy orchestration with SLURM and Kubernetes, optimize interconnects, and provide 24/7 on‑call support for critical deployments, collaborating with top engineering teams to push frontier AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Senior GPU Infrastructure Architect for Frontier AI
Senior GPU Infrastructure Architect for Frontier AI

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Equity incentives
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior GPU Cloud Operations Engineer
Senior GPU Cloud Operations Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior GPU HPC Infra Engineer for AI Training
Senior GPU HPC Infra Engineer for AI Training

Showcify • United States

On-site
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
RRSP/401K
+4
Senior GPU Compute Cluster Architect
Senior GPU Compute Cluster Architect

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Equity incentives