Lead GPU Cluster Solution Architect

Axe Compute

Miami (FL)

On-site

USD 140,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Axe Compute in Miami is seeking a Lead Architect, GPU Cluster Solutions to craft buildable GPU cluster designs aligned with NVIDIA Reference Architecture for signed engagements. You will translate client requirements into a complete design covering compute, storage, and networking.

Lead the sparing, availability, and site-constraint adaptations, collaborating with deployments leadership to ensure the design translates into actionable project plans and validated acceptance criteria.

Qualifications

  • 7+ years in solutions architecture or systems engineering for GPU/HPC infrastructure.
  • Deep knowledge of NVIDIA Reference Architecture (HGX, NVL72).
  • Hands-on experience with InfiniBand, RoCE, and high-speed Ethernet.
  • Experience incorporating firewall, VPN, and dedicated circuit requirements into designs.
  • Experience with high-speed shared storage solutions.

Responsibilities

  • Design GPU cluster configurations against NVIDIA Reference Architecture for each signed engagement.
  • Translate client requirements into a complete bill of design with compute, storage, and networking.
  • Design network topology and fabric selection (InfiniBand/RoCE/Ethernet) per workload.
  • Incorporate internet/VPN/firewall connectivity into cluster designs.
  • Formulate hot/cold sparing and availability plans for SLAs.
  • Adapt designs to site power, cooling and space constraints with procurement.
  • Collaborate with VP Deployments to translate designs into buildable project plans.
  • Support acceptance test design and criteria for as-designed architecture.

Skills

GPU HPC infrastructure
Solutions architecture
Network engineering
Data center design
Cross-functional collaboration

Tools

InfiniBand
RoCE
Ethernet fabric design
Firewall/VPN/dedicated circuit
NVIDIA Reference Architecture (HGX/NVL72)

Job description

Axe Compute is seeking a Lead Architect, GPU Cluster Solutions to design GPU cluster configurations for prospective and signed engagements, translating client requirements and NVIDIA Reference Architecture into a buildable, supportable design, spanning compute, storage, networking, software and spares strategy for support.

ROLE AT A GLANCE
  • Mandate: Own end-to-end technical design of GPU cluster deployments, from client requirements through NVIDIA Reference Architecture compliance and sparing strategy.
  • Scope: Cluster design, network architecture (InfiniBand/RoCE/Ethernet), sparing and spares planning, connectivity design (internet/VPN/firewall, dedicated circuits), design adjustments for site and hardware constraints.
  • Key Outcomes: Designs that meet client SLAs and NVIDIA Reference Architecture standards, sparing plans that protect uptime commitments, designs that account for real-world site and hardware lead-time constraints.
WHAT YOU WILL OWN
Cluster Design & Reference Architecture
  • Design GPU cluster configurations (compute, storage, networking) against NVIDIA Reference Architecture for each signed engagement.
  • Translate client technical requirements into a complete bill of design, including all necessary compute, storage, and networking components.
  • Design network topology and fabric selection, including InfiniBand, RoCE, and Ethernet options, appropriate to each client's workload and performance requirements.
  • Incorporate internet, VPN, and firewall connectivity requirements into cluster designs.
  • Design dedicated point-to-point network requirements where needed, including protected optical circuits and similar dedicated connectivity.
Sparing & Availability Strategy
  • Formulate and own the hot/cold sparing plan for each deployment to meet contracted SLA commitments.
  • Adjust sparing and design assumptions based on data center power/cooling parameters and hardware lead-time constraints.
Design Adaptation & Site Constraints
  • Adjust cluster designs to fit site-specific power, cooling, and space constraints identified by the Data Center Procurement and Operations Director.
  • Work with Supply Chain to align design decisions with realistic hardware delivery timing.
Cross-Functional Collaboration
  • Partner with the VP, Deployments and Deployment Program Manager to ensure designs translate cleanly into buildable, trackable project plans.
  • Support acceptance test design and criteria definition, ensuring test procedures validate the as-designed architecture.
REQUIRED QUALIFICATIONS
  • 7+ years in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure.
  • Deep working knowledge of NVIDIA Reference Architecture (HGX, NVL72) and GPU cluster design principles.
  • Hands-on experience with InfiniBand, RoCE, and high-speed Ethernet fabric design.
  • Experience incorporating firewall, VPN, and dedicated circuit (e.g., protected optical) requirements into network designs.
  • Experience with high-speed shared storage solutions (e.g., Weka, Vast, DDN).
PREFERRED QUALIFICATIONS
  • Experience designing clusters for large enterprise clients or neoclouds, not just internal infrastructure.
  • Familiarity with NVIDIA NCP program requirements and certification processes.
  • Experience with capacity or sparing modeling tools.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Architect – Solutions Lead
Senior GPU Cluster Architect – Solutions Lead

Axe Compute • Miami (FL)

On-site
USD 140,000 - 170,000
Deployment Program Manager
Deployment Program Manager

Axe Compute • Miami (FL)

On-site
USD 120,000 - 180,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
HPC Solution Architect
HPC Solution Architect

Coda Search│Staffing • Dallas (TX)

Hybrid
USD 120,000 - 160,000
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)

NJF Global Holdings Ltd • New York (NY)

On-site
USD 150,000 - 200,000
Lead AI Infrastructure Solutions Architect for HPC Clusters
Lead AI Infrastructure Solutions Architect for HPC Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior Network Solution Architect – AI Fabrics
Senior Network Solution Architect – AI Fabrics

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Senior GPU HPC Cluster Engineer — Equity Eligible
Senior GPU HPC Cluster Engineer — Equity Eligible

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior Network Solutions Architect
Senior Network Solutions Architect

Hamilton Barnes ? • San Francisco (CA)

On-site
USD 190,000 - 260,000
Senior Solutions Architect, Cluster Design and Architecture - Networking
Senior Solutions Architect, Cluster Design and Architecture - Networking

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits