Remote - Lead GPU Cluster Solutions Architect

Orion Placement

United States

Remote

USD 150,000 - 210,000

Full time

12 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Bonus
Equity

Job summary

Orion Placement is seeking an experienced Solutions Architect to own end-to-end GPU cluster deployments, translating client requirements into deployable, scalable architectures. The role emphasizes HGX/NVL72-based designs, high-performance networks, and collaboration with data centers, supply chain, and leadership.

You will work on enterprise-grade AI infrastructure projects, with remote nationwide placement and travel to partner sites as needed.

Qualifications

  • 7+ years of experience in solutions architecture, network engineering, or systems engineering for GPU/HPC infrastructure.
  • Deep working knowledge of NVIDIA Reference Architecture and GPU cluster design principles.
  • Hands-on experience designing InfiniBand, RoCE, and/or high-speed Ethernet fabrics.
  • Proven track record designing GPU or HPC clusters, not just cloud infra.
  • Experience developing sparing and spare strategies for mission-critical infra.
  • Experience with firewalls, VPNs, dedicated circuits, protected optical connectivity in infra designs.

Responsibilities

  • Own end-to-end technical architecture for GPU cluster deployments from customer requirements to deployment-ready design.
  • Design GPU cluster configurations spanning compute, storage, networking, software needs, and supporting infra.
  • Translate client requirements into complete bills of design covering compute, storage, networking, and connectivity components.
  • Apply NVIDIA Reference Architecture principles including HGX and NVL72-based designs.
  • Design high-performance networks using InfiniBand, RoCE, and high-speed Ethernet based on workloads and performance.

Skills

GPU cluster design
InfiniBand networking
RoCE networking
NVIDIA Reference Architecture
HPC infrastructure

Tools

InfiniBand Fabric
RoCE Fabric
NVIDIA Reference Architecture

Job description

  • Own the technical architecture behind next-generation GPU and AI compute deployments.
  • Design sophisticated GPU clusters from client requirements through production-ready architecture.
  • Work directly with NVIDIA Reference Architecture, high-speed networking, storage, connectivity, and availability strategy.
  • Solve challenging infrastructure problems where performance, reliability, power, cooling, and hardware constraints all matter.
  • Have significant technical ownership over designs supporting enterprise and neocloud deployments.
  • Collaborate closely with deployment, data center, supply chain, and program leadership to turn architecture into real-world infrastructure.
  • Join a fast-growing AI infrastructure environment where your technical decisions directly impact customer outcomes.
  • Competitive bonus and equity opportunity in addition to base compensation.

Location: Remote nationwide, with preference for candidates based in the US. Travel to data center partner sites is required.

Note: Must have 7+ years of directly relevant experience in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure. Candidates must have hands‑on GPU cluster design experience plus strong InfiniBand, RoCE, or high‑speed Ethernet networking expertise. Generic cloud architecture or enterprise networking experience without meaningful GPU/HPC infrastructure exposure will not meet the requirements.

About Us

We are building next-generation AI infrastructure that gives enterprises access to high-performance GPU compute with speed, flexibility, and reliability. Our technical teams design and deploy sophisticated GPU clusters across data center environments, and we are looking for an architect who can turn demanding customer requirements into robust, buildable infrastructure. Confidential Employer.

Job Description
  • Own end‑to‑end technical architecture for GPU cluster deployments from customer requirements through deployment‑ready design.
  • Design GPU cluster configurations spanning compute, storage, networking, software requirements, and supporting infrastructure.
  • Translate client technical requirements into complete bills of design covering all required compute, storage, networking, and connectivity components.
  • Apply NVIDIA Reference Architecture principles, including HGX and NVL72-based GPU cluster designs.
  • Design high‑performance network fabrics using InfiniBand, RoCE, and high‑speed Ethernet based on workload and performance requirements.
  • Incorporate internet, VPN, firewall, dedicated circuit, protected optical, and other connectivity requirements into cluster architectures.
  • Develop hot and cold sparing strategies designed to meet contracted availability and SLA commitments.
  • Adapt cluster designs to site‑specific power, cooling, space, hardware, and deployment constraints.
  • Partner with data center teams to account for real‑world facility limitations when finalizing technical architecture.
  • Work with Supply Chain to ensure architecture decisions align with realistic hardware availability and lead times.
  • Partner with deployment leadership and program management to translate designs into executable build plans.
  • Support acceptance test planning and define technical criteria that validate the deployed architecture against the approved design.
  • Evaluate and incorporate high‑speed shared storage solutions such as Weka, VAST Data, and DDN where appropriate.
  • Maintain technical ownership of architecture decisions while balancing performance, availability, cost, schedule, and operational supportability.
Qualifications
  • 7+ years of experience in solutions architecture, network engineering, systems engineering, or similar roles supporting GPU, HPC, or large‑scale compute infrastructure.
  • Deep working knowledge of NVIDIA Reference Architecture and GPU cluster design principles.
  • Hands‑on experience designing InfiniBand, RoCE, and/or high‑speed Ethernet fabrics.
  • Proven experience designing GPU or HPC clusters rather than solely consuming cloud infrastructure.
  • Experience developing sparing and spares strategies for mission‑critical infrastructure.
  • Experience integrating firewalls, VPNs, dedicated circuits, protected optical connectivity, and related networking requirements into infrastructure designs.
  • Experience with high‑speed shared storage technologies such as Weka, VAST Data, or DDN.
  • Strong understanding of compute, storage, networking, and data center infrastructure dependencies.
  • Ability to translate complex customer requirements into complete, practical, buildable technical architectures.
  • Strong cross‑functional communication and documentation skills.
  • Experience supporting enterprise customers or neocloud deployments is preferred.
  • NVIDIA NCP program or certification experience is a plus.
  • Experience with capacity planning or sparing modeling tools is a plus.
Why You Will Love Working Here
  • Work on technically challenging GPU infrastructure projects at the center of the AI compute market.
  • Own architecture decisions that directly influence performance, reliability, scalability, and customer success.
  • Gain exposure to cutting‑edge NVIDIA GPU architectures and high‑speed networking technologies.
  • Collaborate with experienced infrastructure, deployment, data center, supply chain, and executive teams.
  • Work remotely while remaining closely connected to real‑world data center deployments.
  • Opportunity to help establish repeatable architecture standards as the business scales.
  • Competitive base compensation plus bonus and equity.
  • Make a visible impact in a high‑growth environment where strong technical judgment is valued.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Cluster Architect
GPU Cluster Architect

Jobgether SRL • United States

Remote
USD 184,000 - 318,000
Medical, dental, vision insurance
Remote work reimbursement
RSUs may be available
+3
Principal Hardware Design Engineer
Principal Hardware Design Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 180,000 - 230,000
Medical, dental, and vision insurance
401(k) plan with company match
Paid holidays
Applied Researcher – Network Expert
Applied Researcher – Network Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 120,000 - 190,000
Solutions Architect - NVIDIA Cloud Partners
Solutions Architect - NVIDIA Cloud Partners

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity compensation
Benefits
Solutions Architect - NVIDIA Cloud Partners
Solutions Architect - NVIDIA Cloud Partners

NVIDIA • Virginia (IL)

On-site
USD 184,000 - 357,000
Senior Solutions Architect, Generative AI
Senior Solutions Architect, Generative AI

NVIDIA Corporation • Santa Clara (CA)

Remote
USD 184,000 - 357,000
Equity
Benefits
Solutions Architect, HPC Systems Engineer
Solutions Architect, HPC Systems Engineer

NVIDIA Corporation • Santa Clara (CA)

Remote
USD 184,000 - 357,000
Equity
Benefits
Hardware Design Engineer
Hardware Design Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Manager, Solutions Architecture – GPU and Networking Systems
Manager, Solutions Architecture – GPU and Networking Systems

NVIDIA • Santa Clara (CA)

Remote
USD 224,000 - 431,000
Equity
Benefits
Remote work option
Senior Solutions Architect - AI Infrastructure
Senior Solutions Architect - AI Infrastructure

NVIDIA • California (MO)

On-site
USD 184,000 - 356,500
Equity
Benefits