Senior GPU Compute Cluster Architect

Blue Signal Search

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Blue Signal Search is seeking a highly technical infrastructure expert to lead the design, deployment, and optimization of large-scale GPU computing environments in San Francisco. You will work with an experienced engineering team building from the ground up on cutting-edge AI workloads.

This role emphasizes ownership and deep technical focus, with collaboration directly with customers to tailor optimized infrastructure solutions and drive performance across the full GPU stack.

Qualifications

  • Proven hands-on experience designing, deploying, and operating production-scale GPU clusters.
  • Deep expertise implementing InfiniBand and/or RoCEv2 networking fabrics.
  • Strong experience with CUDA and/or ROCm environments including firmware, drivers, NCCL/RCCL, and performance analysis.
  • Demonstrated success supporting customer-facing production infrastructure.
  • Excellent troubleshooting across compute, networking, storage, and distributed systems.
  • Ability to explain complex technical concepts to engineering teams and customers.
  • Preference for individual contributor role with technical excellence.

Responsibilities

  • Architect and deploy high-performance GPU compute clusters for production AI workloads.
  • Design, configure, and optimize InfiniBand/RoCEv2 networking environments.
  • Drive performance tuning across the GPU software stack (CUDA, ROCm, NCCL/RCCL, drivers).
  • Troubleshoot infrastructure, networking, and software bottlenecks in large-scale deployments.
  • Develop scalable infrastructure standards, deployment methodologies, and best practices.
  • Collaborate with customers to understand workload requirements and deliver optimized solutions.
  • Lead technical investigations during production incidents and guide resolution.
  • Create operational docs, procedures, and performance validation for future deployments.

Skills

GPU clusters
InfiniBand/RoCEv2
CUDA/ROCm
Performance analysis
Troubleshooting
Customer-facing
Distributed systems
Documentation

Job description

Blue Signal Search is seeking a highly technical infrastructure expert to lead the design, deployment, and optimization of large-scale GPU computing environments in San Francisco. You will work with an experienced engineering team building from the ground up on cutting-edge AI workloads.

This role emphasizes ownership and deep technical focus, with collaboration directly with customers to tailor optimized infrastructure solutions and drive performance across the full GPU stack.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Head of GPU Systems Engineering
Head of GPU Systems Engineering

Blue Signal Search • Fremont (CA)

On-site
USD 180,000 - 260,000
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Hybrid GPU Data Center Engineer: Automation & AI Infra
Hybrid GPU Data Center Engineer: Automation & AI Infra

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Senior GPU Cluster Architect for AI Infra at Scale
Senior GPU Cluster Architect for AI Infra at Scale

Partner Company • United States

Remote
USD 184,000 - 318,000
Medical insurance
Dental insurance
Vision insurance
+1
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
GPU Data Center Engineer — Hybrid/Remote
GPU Data Center Engineer — Hybrid/Remote

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2