Senior Network Solutions Architect

Hamilton Barnes ?

San Francisco (CA)

On-site

USD 190,000 - 260,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Stealth Neocloud in San Francisco is seeking a Senior Network Solutions Architect to own the design, deployment, and performance of its GPU compute clusters end to end. You will be hands-on, building the fabric from the rack up and serving as the technical authority, with direct access to founders and no bureaucracy.

This founding-level role requires onsite work and relocation if needed, with opportunities as the team expands.

Qualifications

  • Hands-on experience standing up GPU clusters at production scale, with clusters you built.
  • Real InfiniBand or RoCEv2 design and implementation experience.
  • Comfortable across CUDA and ROCm ecosystems, or able to ramp on a new GPU stack quickly.
  • Direct customer-facing technical experience, the go-to person when something goes wrong.
  • Founding-level ownership and visibility; onsite in San Francisco.

Responsibilities

  • Design compute, storage, and networking topology for new cluster deployments.
  • Specify node configurations, redundancy, and scaling plans.
  • Make real-time build decisions on-site during deployment, not just on paper.
  • Bridge datacenter specifications to cluster deployment, including power whips and rack fit‑out.
  • Design and personally implement GPU interconnect fabric using InfiniBand and/or RoCEv2.
  • Plan and validate bandwidth, topology, and east‑west throughput at scale.
  • Diagnose and fix network bottlenecks under load, hands‑on rather than in theory.
  • Get workloads running well across both CUDA and ROCm stacks.
  • Own driver and firmware compatibility, NCCL/RCCL tuning, and performance benchmarking.
  • Troubleshoot low‑level issues directly across drivers, firmware, and fabric managers.
  • Lead technical onboarding and workload validation for new customers.
  • Act as the direct escalation point when a customer's cluster underperforms, diagnosing it yourself.
  • Translate customer workload requirements into concrete infrastructure and configuration decisions.
  • Document runbooks and SOPs that reflect how the infrastructure actually gets built and fixed.

Skills

Hands-on GPU clusters
InfiniBand RoCEv2
CUDA ROCm ecosystems
Customer-facing engineering
Hands-on leadership
San Francisco onsite
Founding ownership

Tools

InfiniBand hardware
RoCEv2
CUDA
ROCm

Job description

Job Title: Senior Network Solutions Architect

Role:

Join a Stealth Neocloud building GPU compute infrastructure from the ground up in San Francisco. As one of the founding team, the systems you design and build this year are the systems the company runs on. You'll have direct access to and collaboration with the founders, high ownership, high visibility, and no bureaucracy standing between you and the rack.

As a Senior Network Solutions Architect, you will own the design, deployment, and performance of the company's GPU compute clusters end to end, from the network fabric up through the software stack; combined with a genuinely customer-facing role. This is emphatically not a management position. You will be the person racking, cabling, configuring, benchmarking, and debugging the infrastructure yourself, not delegating it, while also leading technical onboarding and acting as the direct escalation point for customers (Obviously you wouldn't be doing physical work every day, but it's a start-up...).

You’ll operate as the technical authority on cluster architecture, fabric engineering, and multi-vendor GPU enablement, reporting directly to the founders and making real-time build decisions on-site during deployments. As the team grows, you may build out a small team under you, but the expectation is that you stay hands‑on and technical rather than shifting into a managerial role.

Responsibilities:
  • Design compute, storage, and networking topology for new cluster deployments
  • Specify node configurations, redundancy, and scaling plans
  • Make real-time build decisions on-site during deployment, not just on paper
  • Bridge datacenter specifications to cluster deployment, including power whips and rack fit‑out
  • Design and personally implement GPU interconnect fabric using InfiniBand and/or RoCEv2
  • Plan and validate bandwidth, topology, and east‑west throughput at scale
  • Diagnose and fix network bottlenecks under load, hands‑on rather than in theory
  • Get workloads running well across both CUDA and ROCm stacks
  • Own driver and firmware compatibility, NCCL/RCCL tuning, and performance benchmarking
  • Troubleshoot low‑level issues directly across drivers, firmware, and fabric managers
  • Lead technical onboarding and workload validation for new customers
  • Act as the direct escalation point when a customer's cluster underperforms, diagnosing it yourself
  • Translate customer workload requirements into concrete infrastructure and configuration decisions
  • Document runbooks and SOPs that reflect how the infrastructure actually gets built and fixed
Skills/Must have:
  • Direct, hands‑on experience standing up GPU clusters at production scale, with specific clusters you built rather than systems you oversaw.
  • Real InfiniBand or RoCEv2 design and implementation experience.
  • Comfortable working across both CUDA and ROCm ecosystems, or clearly demonstrated ability to ramp fast on a new GPU vendor stack.
  • Direct customer‑facing technical experience, comfortable being the person a customer talks to when something's wrong.
  • Genuine preference for staying hands‑on over moving into pure management.
  • Based in or willing to relocate to San Francisco (onsite).
  • Founding‑level ownership and visibility.
  • Direct access to and collaboration with the founders.
  • Onsite role in San Francisco and relocation provided if needed.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Solution Architect - AI Infrastructure
Senior Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
Founding GPU Cluster Architect — Hands-On, SF
Founding GPU Cluster Architect — Hands-On, SF

Hamilton Barnes ? • San Francisco (CA)

On-site
USD 190,000 - 260,000
Senior HPC Engineer
Senior HPC Engineer

Hamilton Barnes ? • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation provided
Founding-level ownership
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
HPC/ GPU Cluster Architect
HPC/ GPU Cluster Architect

Electric Capital • San Francisco (CA)

Hybrid
USD 220,000 - 300,000
Generous equity grant
401(k) matching
Comprehensive medical, dental, and vision insurance
+3
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior GPU Compute Infra Architect - Onsite SF
Senior GPU Compute Infra Architect - Onsite SF

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
HPC/ GPU Hardware Engineer
HPC/ GPU Hardware Engineer

The San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Generous equity grant
Retirement matching
Comprehensive medical, dental, and vision insurance
+5
GPU Infrastructure Engineer
GPU Infrastructure Engineer

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
Senior Solutions Architect, Cluster Design and Architecture - Networking
Senior Solutions Architect, Cluster Design and Architecture - Networking

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits