Cluster Design

Blue Signal Search

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Blue Signal Search is seeking a highly technical infrastructure expert to lead the design, deployment, and optimization of large-scale GPU computing environments in San Francisco. You will work with an experienced engineering team building from the ground up on cutting-edge AI workloads.

This role emphasizes ownership and deep technical focus, with collaboration directly with customers to tailor optimized infrastructure solutions and drive performance across the full GPU stack.

Qualifications

  • Proven hands-on experience designing, deploying, and operating production-scale GPU clusters.
  • Deep expertise implementing InfiniBand and/or RoCEv2 networking fabrics.
  • Strong experience with CUDA and/or ROCm environments including firmware, drivers, NCCL/RCCL, and performance analysis.
  • Demonstrated success supporting customer-facing production infrastructure.
  • Excellent troubleshooting across compute, networking, storage, and distributed systems.
  • Ability to explain complex technical concepts to engineering teams and customers.
  • Preference for individual contributor role with technical excellence.

Responsibilities

  • Architect and deploy high-performance GPU compute clusters for production AI workloads.
  • Design, configure, and optimize InfiniBand/RoCEv2 networking environments.
  • Drive performance tuning across the GPU software stack (CUDA, ROCm, NCCL/RCCL, drivers).
  • Troubleshoot infrastructure, networking, and software bottlenecks in large-scale deployments.
  • Develop scalable infrastructure standards, deployment methodologies, and best practices.
  • Collaborate with customers to understand workload requirements and deliver optimized solutions.
  • Lead technical investigations during production incidents and guide resolution.
  • Create operational docs, procedures, and performance validation for future deployments.

Skills

GPU clusters
InfiniBand/RoCEv2
CUDA/ROCm
Performance analysis
Troubleshooting
Customer-facing
Distributed systems
Documentation

Job description

A rapidly growing AI infrastructure company is seeking a highly technical infrastructure expert to lead the design, deployment, and optimization of large-scale GPU computing environments. This is an opportunity to have a direct impact on the architecture powering advanced AI workloads while working alongside an experienced engineering team building from the ground up.

If you thrive on solving complex infrastructure challenges, enjoy working directly with cutting‑edge hardware, and prefer remaining deeply technical rather than moving into management, this role offers exceptional ownership and influence.

What You'll Do
  • Architect and deploy high performance GPU compute clusters supporting production AI and machine learning workloads.
  • Design, configure, and optimize high speed networking environments using InfiniBand and or RoCEv2 technologies.
  • Drive performance tuning across the complete GPU software stack, including CUDA, ROCm, drivers, firmware, NCCL, RCCL, and system benchmarking.
  • Troubleshoot infrastructure, networking, and software bottlenecks affecting large scale distributed compute environments.
  • Develop scalable infrastructure standards, deployment methodologies, and operational best practices.
  • Collaborate directly with customers to understand workload requirements and deliver optimized infrastructure solutions.
  • Lead technical investigations during production incidents and provide expert guidance through issue resolution.
  • Create operational documentation, implementation procedures, and performance validation processes for future deployments.
Required Qualifications
  • Proven hands‑on experience designing, deploying, and operating production‑scale GPU clusters.
  • Deep expertise implementing InfiniBand and or RoCEv2 networking fabrics.
  • Strong experience working with CUDA and or ROCm environments, including firmware management, driver integration, NCCL or RCCL optimization, and performance analysis.
  • Demonstrated success supporting customer facing production infrastructure.
  • Excellent troubleshooting skills spanning compute, networking, storage, and distributed systems.
  • Strong communication skills with the ability to explain complex technical concepts to both engineering teams and customers.
  • Preference for remaining an individual contributor focused on technical excellence rather than pursuing people management responsibilities.
Preferred Experience
  • Background supporting AI, HPC, or accelerated computing environments.
  • Experience evaluating system performance through benchmarking and workload optimization.
  • Familiarity with distributed infrastructure operations and production support methodologies.
  • Passion for building reliable, scalable infrastructure that directly impacts customer success.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Senior Solution Architect - AI Infrastructure
Senior Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000