Senior Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 233,000 - 316,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco

Job summary

Hamilton Barnes Associates Limited is seeking a Senior Solutions Architect / Network Infrastructure Engineer to design, deploy, and optimize GPU compute clusters. You will work directly with the founders in a high-ownership, hands-on environment, shaping the platform’s core infrastructure onsite in San Francisco.

You will lead topology design, hardware configurations, and bottleneck troubleshooting, bridging data-center specs to deployment while staying hands-on and customer-focused.

Qualifications

  • Hands-on experience standing up GPU clusters at production scale.
  • Experience with InfiniBand or RoCEv2 design and deployment.
  • Ability to work directly with customers and founders.
  • Willing to be onsite in San Francisco.

Responsibilities

  • Design compute, storage, and networking topology for new cluster deployments.
  • Specify node configurations, redundancy, and scaling plans.
  • Lead deployment and troubleshooting on-site.
  • Bridge data center specs to cluster deployment and rack fit-out.
  • Implement GPU interconnect fabric using InfiniBand and/or RoCEv2.
  • Document runbooks and SOPs reflecting how the infrastructure actually gets built and fixed.

Skills

GPU cluster deployment
InfiniBand design
RoCEv2 implementation
CUDA ROCm expertise
Customer-facing engineering
Hands-on engineer

Tools

InfiniBand hardware
RoCE fabric

Job description

Looking for a role with plenty of growth opportunities?

Join an early-stage company building GPU compute infrastructure from the ground up. As one of the first technical hires, you'll help shape the company's core infrastructure while working directly with the founders in a high-ownership, hands-on environment.

This company is seeking a Senior Solutions Architect / Network Infrastructure Engineer to lead the design, deployment, and optimisation of GPU compute clusters across networking, software, and infrastructure. This customer-facing role combines hands-on engineering with technical leadership, including cluster architecture, fabric engineering, deployment, troubleshooting, and customer onboarding, while influencing the future direction of the platform.

Responsibilities:
  • Design compute, storage, and networking topology for new cluster deployments
  • Specify node configurations, redundancy, and scaling plans
  • Make real-time build decisions on-site during deployment, not just on paper
  • Bridge data center specifications to cluster deployment, including power whips and rack fit-out
  • Design and personally implement GPU interconnect fabric using InfiniBand and/or RoCEv2
  • Plan and validate bandwidth, topology, and east-west throughput at scale
  • Diagnose and fix network bottlenecks under load, hands-on rather than in theory
  • Get workloads running well across both CUDA and ROCm stacks
  • Own driver and firmware compatibility, NCCL/RCCL tuning, and performance benchmarking
  • Troubleshoot low-level issues directly across drivers, firmware, and fabric managers
  • Lead technical onboarding and workload validation for new customers
  • Act as the direct escalation point when a customer's cluster underperforms, diagnosing it yourself
  • Translate customer workload requirements into concrete infrastructure and configuration decisions
  • Document runbooks and SOPs that reflect how the infrastructure actually gets built and fixed
Skills/Must Have:
  • Direct, hands-on experience standing up GPU clusters at production scale, with specific clusters you built rather than systems you oversaw
  • Real InfiniBand or RoCEv2 design and implementation experience
  • Comfortable working across both CUDA and ROCm ecosystems, or clearly demonstrated ability to ramp fast on a new GPU vendor stack
  • Direct customer-facing technical experience, comfortable being the person a customer talks to when something's wrong
  • Genuine preference for staying hands-on over moving into pure management
  • Based in or willing to relocate to San Francisco (onsite)
Benefits:
  • Founding-level ownership and visibility
  • Direct access to and collaboration with the founders
  • Onsite role in San Francisco
Salary:
  • $275,000 Base Salary
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Customer Solution Architect - Systems Integrator
Customer Solution Architect - Systems Integrator

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 225,000 - 275,000
RSU equity
20% bonus
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior AI Network Engineer - AI Infrastructure
Senior AI Network Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
HPC Engineer - AI Infrastructure
HPC Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Senior Solutions Architect, AI Infrastructure, Senior Solutions Architect, AI Infrastructure
Senior Solutions Architect, AI Infrastructure, Senior Solutions Architect, AI Infrastructure

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equities and benefits
Remote work options
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000