Senior GPU Infrastructure Architect

Primeintellect

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Prime Intellect in San Francisco is seeking a senior infrastructure engineer to design and deploy GPU-heavy workloads at scale. You will work directly with customers to translate workload requirements into robust GPU clusters and deployment strategies.

The role combines hands-on engineering with high-visibility architecture discussions, requiring deep knowledge of HPC tooling, networking, and production reliability.

Qualifications

  • 3+ years hands-on experience with GPU clusters and HPC environments.
  • Deep expertise with SLURM and Kubernetes in production GPU settings.
  • Proven experience with InfiniBand configuration and troubleshooting.
  • Strong understanding of NVIDIA GPU architecture, CUDA ecosystem, and driver stack.
  • Experience with infrastructure automation tools (Ansible, Terraform).
  • Proficiency in Python, Bash, and systems programming.
  • Track record of customer-facing technical leadership.

Responsibilities

  • Partner with clients to understand workload requirements and design optimal GPU cluster architectures.
  • Create technical proposals and capacity planning for clusters ranging from 100 to 10,000+ GPUs.
  • Develop deployment strategies for LLM training, inference, and HPC workloads.
  • Present architectural recommendations to technical and executive stakeholders.
  • Deploy and configure orchestration systems including SLURM and Kubernetes for distributed workloads.
  • Implement high‑performance networking with InfiniBand, RoCE, and NVLink interconnects.
  • Optimize GPU utilization, memory management, and inter‑node communication.
  • Configure parallel filesystems (Lustre, BeeGFS, GPFS) for optimal I/O performance.
  • Tune system performance from kernel parameters to CUDA configurations.
  • Serve as primary technical escalation point for customer infrastructure issues.
  • Diagnose and resolve complex problems across the full stack - hardware, drivers, networking, and software.
  • Implement monitoring, alerting, and automated remediation systems.
  • Provide 24/7 on‑call support for critical customer deployments.
  • Create runbooks and documentation for customer operations teams.

Skills

GPU clusters and HPC environments
customer-facing technical leadership
Python
Bash
systems programming
CUDA ecosystem
driver stack

Tools

SLURM
Kubernetes
InfiniBand
NVIDIA CUDA
Docker
Containerd
Enroot
Terraform
Ansible

Job description

Prime Intellect in San Francisco is seeking a senior infrastructure engineer to design and deploy GPU-heavy workloads at scale. You will work directly with customers to translate workload requirements into robust GPU clusters and deployment strategies.

The role combines hands-on engineering with high-visibility architecture discussions, requiring deep knowledge of HPC tooling, networking, and production reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

sfcompute • San Francisco (CA)

On-site
USD 180,000 - 240,000
Generous equity grant
Visa Sponsorships
Retirement matching
+5
Senior Datacenter Networking Engineer (GPU Infra)
Senior Datacenter Networking Engineer (GPU Infra)

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Infrastructure Operations Lead
GPU Infrastructure Operations Lead

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cloud Operations Engineer
Senior GPU Cloud Operations Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior GPU Compute Cluster Architect
Senior GPU Compute Cluster Architect

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Staff Network Engineer — GPU Data Center & HPC Networking
Staff Network Engineer — GPU Data Center & HPC Networking

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2