Senior GPU Compute Cluster Architect

Blue Signal Search

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Blue Signal Search is seeking a highly technical infrastructure expert to lead the design, deployment, and optimization of large-scale GPU computing environments in San Francisco. You will work with an experienced engineering team building from the ground up on cutting-edge AI workloads.

This role emphasizes ownership and deep technical focus, with collaboration directly with customers to tailor optimized infrastructure solutions and drive performance across the full GPU stack.

Qualifications

  • Proven hands-on experience designing, deploying, and operating production-scale GPU clusters.
  • Deep expertise implementing InfiniBand and/or RoCEv2 networking fabrics.
  • Strong experience with CUDA and/or ROCm environments including firmware, drivers, NCCL/RCCL, and performance analysis.
  • Demonstrated success supporting customer-facing production infrastructure.
  • Excellent troubleshooting across compute, networking, storage, and distributed systems.
  • Ability to explain complex technical concepts to engineering teams and customers.
  • Preference for individual contributor role with technical excellence.

Responsibilities

  • Architect and deploy high-performance GPU compute clusters for production AI workloads.
  • Design, configure, and optimize InfiniBand/RoCEv2 networking environments.
  • Drive performance tuning across the GPU software stack (CUDA, ROCm, NCCL/RCCL, drivers).
  • Troubleshoot infrastructure, networking, and software bottlenecks in large-scale deployments.
  • Develop scalable infrastructure standards, deployment methodologies, and best practices.
  • Collaborate with customers to understand workload requirements and deliver optimized solutions.
  • Lead technical investigations during production incidents and guide resolution.
  • Create operational docs, procedures, and performance validation for future deployments.

Skills

GPU clusters
InfiniBand/RoCEv2
CUDA/ROCm
Performance analysis
Troubleshooting
Customer-facing
Distributed systems
Documentation

Job description

Blue Signal Search is seeking a highly technical infrastructure expert to lead the design, deployment, and optimization of large-scale GPU computing environments in San Francisco. You will work with an experienced engineering team building from the ground up on cutting-edge AI workloads.

This role emphasizes ownership and deep technical focus, with collaboration directly with customers to tailor optimized infrastructure solutions and drive performance across the full GPU stack.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior GPU Compute Infra Architect - Onsite SF
Senior GPU Compute Infra Architect - Onsite SF

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
Hybrid GPU Data Center Engineer: Automation & AI Infra
Hybrid GPU Data Center Engineer: Automation & AI Infra

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
GPU Data Center Engineer — Hybrid/Remote
GPU Data Center Engineer — Hybrid/Remote

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior MLOps Engineer: GPU AI Infra & Production
Senior MLOps Engineer: GPU AI Infra & Production

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
AI Kernel & GPU HPC Cluster Engineer
AI Kernel & GPU HPC Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
GPU Infra Solutions Architect for Large-Scale AI Clusters
GPU Infra Solutions Architect for Large-Scale AI Clusters

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000