Senior GPU Infrastructure Architect for Frontier AI

Prime Intellect

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 300,000

Full time

24 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity incentives

Job summary

Prime Intellect is building the open superintelligence stack, delivering a full-stack platform for post-training at frontier scale. You will design GPU cluster architectures, plan capacity for large-scale deployments, and develop deployment strategies for LLM training and HPC workloads.

You will optimize networking, file systems, and driver stacks while maintaining 24/7 support for critical customer deployments.

Qualifications

  • 3+ years hands-on experience with GPU clusters and HPC environments.
  • Deep expertise with SLURM and Kubernetes in production GPU settings.
  • Proven experience with InfiniBand configuration and troubleshooting.
  • Strong understanding of NVIDIA GPU architecture, CUDA ecosystem, and driver stack.
  • Experience with infrastructure automation tools (Ansible, Terraform).
  • Proficiency in Python, Bash, and systems programming.
  • Track record of customer-facing technical leadership.

Responsibilities

  • Partner with clients to understand workload requirements and design optimal GPU cluster architectures.
  • Create technical proposals and capacity planning for clusters ranging from 100 to 10,000+ GPUs.
  • Develop deployment strategies for LLM training, inference, and HPC workloads.
  • Present architectural recommendations to technical and executive stakeholders.
  • Deploy and configure orchestration systems including SLURM and Kubernetes for distributed workloads.
  • Implement high-performance networking with InfiniBand, RoCE, and NVLink interconnects.
  • Optimize GPU utilization, memory management, and inter-node communication.
  • Configure parallel filesystems for optimal I/O performance.
  • Tune system performance from kernel parameters to CUDA configurations.
  • Serve as primary technical escalation point for customer infrastructure issues.
  • Diagnose and resolve complex problems across the full stack.
  • Implement monitoring, alerting, and automated remediation systems.
  • Provide 24/7 on-call support for critical customer deployments.
  • Create runbooks and documentation for customer operations teams.

Skills

GPU clusters
SLURM
Kubernetes
InfiniBand
CUDA ecosystem
Python
Terraform
Ansible
Customer leadership

Tools

Docker
Containerd
Lustre

Job description

Prime Intellect is building the open superintelligence stack, delivering a full-stack platform for post-training at frontier scale. You will design GPU cluster architectures, plan capacity for large-scale deployments, and develop deployment strategies for LLM training and HPC workloads.

You will optimize networking, file systems, and driver stacks while maintaining 24/7 support for critical customer deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cloud Operations Engineer
Senior GPU Cloud Operations Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Frontier AI GPU Storage Infrastructure Engineer
Frontier AI GPU Storage Infrastructure Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Frontier AI Training Systems Engineer
Frontier AI Training Systems Engineer

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Cash compensation and equity Incentive
Remote or SF office
Visa sponsorship and relocation
+2
Staff Storage Engineer – Frontier AI Infra
Staff Storage Engineer – Frontier AI Infra

Prime-Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Frontier AI Storage Infrastructure Engineer
Frontier AI Storage Infrastructure Engineer

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Storage Systems Engineer for Frontier AI
Storage Systems Engineer for Frontier AI

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Datacenter Networking for GPU AI Infrastructure
Staff Datacenter Networking for GPU AI Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000