HPC Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 270,000 - 330,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Healthcare + Dental + Vision
401(k)
Unlimited PTO

Job summary

Hamilton Barnes Associates Limited is seeking the first engineer in a new function to build and define an AI compute platform from the ground up in San Francisco. The role focuses on evaluating and onboarding GPU compute providers, designing rigorous benchmarks, and establishing technical standards.

The ideal candidate has deep HPC experience with GPU clusters, strong network fabric knowledge, and hands-on HPC software expertise.

Qualifications

  • Deep HPC experience with GPU clusters at meaningful scale.
  • Expertise in InfiniBand/RoCE network fabrics and topology diagnosis.
  • Hands-on with Slurm, Kubernetes, OpenMPI, or equivalent HPC stack.
  • Knowledge of data-center physics: power, cooling, cabling.
  • Benchmarking judgment and design of robust tests for large-scale training.
  • Ability to write technical standards providers can build against.
  • Comfort delivering a failing grade to a provider while preserving relationships.

Responsibilities

  • Vet prospective compute providers across GPU hardware, network, storage, and orchestration — and decide qualification.
  • Build acceptance test suite, benchmark methodology, and quality thresholds from scratch.
  • Execute hands-on validation: burn‑in testing, fabric validation, NCCL benchmarks, storage tests.
  • Guide providers through onboarding, collaborating with their engineers to close gaps on schedule.
  • Partner with compute procurement on technical due diligence and remediation costs before contracts.
  • Maintain standing technical relationships with providers to catch architectural drift.
  • Feed learnings back into provider standards to speed up cycles.

Skills

HPC experience
GPU clusters
InfiniBand RoCE
Slurm
Kubernetes
OpenMPI
Data-centre literacy
Benchmarking
Technical standards
Provider management

Tools

Slurm
Kubernetes
OpenMPI

Job description

Keen to join a company that champions growth and development?

Join a fast-scaling AI compute platform operating at the forefront of global AI infrastructure. The organisation is building the technology that connects AI training and inference workloads with a global network of compute providers, supporting some of the world's leading AI labs and cloud operators.

This company is seeking the first engineer in this function to build and define the role from the ground up. This greenfield opportunity is ideal for someone with deep HPC experience who wants to shape technical standards, establish best practices, and create the foundations of a rapidly growing AI infrastructure platform.

Responsibilities:
  • You will be vetting prospective compute providers across GPU hardware, network fabric, storage, and orchestration — and making the call on whether a cluster qualifies
  • You will be building the acceptance test suite, benchmark methodology, and quality thresholds from scratch, replacing tribal knowledge with rigorous written standards
  • You will be executing hands-on validation where it counts: burn-in testing, fabric validation (InfiniBand/RoCE), NCCL benchmarks, and storage performance testing
  • You will be guiding providers through technical onboarding, working directly with their engineers to close gaps on a predictable timeline
  • You will be partnering with compute procurement on technical due diligence - giving a clear read on remediation cost before contracts are signed
  • You will be the standing technical relationship with existing providers, catching architectural drift before it becomes a customer issue
  • You will be feeding learnings back into provider-facing standards so each qualification cycle is faster than the last
Skills / Must Have:
  • Deep HPC experience, you will have designed, built, or operated GPU clusters at meaningful scale
  • Strong network fabric knowledge across both InfiniBand and RoCE; you will be able to evaluate topology and diagnose underperformance
  • Hands‑on with distributed orchestration and the HPC software stack: Slurm, Kubernetes, OpenMPI or equivalent
  • Data‑centre literacy, power, cooling, cabling, physical-layer realities; you will know what questions to ask when you walk a facility
  • Benchmarking judgement - you will know which numbers matter for large-scale training and how to design tests that can't be gamed
  • You will be writing technical standards that providers can build against without you in the room
  • You will be comfortable delivering a failing grade to a provider who wants your business - and keeping the relationship intact
Benefits:
  • Large equity
  • Comprehensive healthcare, dental, and vision (you and dependents)
  • 401(k)
  • Unlimited PTO
Salary:
  • $300,000 base salary
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
HPC Engineer - AI Infrastructure
HPC Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
Customer Solution Architect - Systems Integrator
Customer Solution Architect - Systems Integrator

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 225,000 - 275,000
RSU equity
20% bonus
Senior AI Network Engineer - AI Infrastructure
Senior AI Network Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
Senior Solution Architect - AI Infrastructure
Senior Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1