AI Cluster Architect — Power-Aware GPU Scaling

Webhosting

Northern (KY)

Hybrid

USD 165,000 - 185,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical Benefits
401(k) Match
PD Reimbursement
PTO & Holidays
Sabbatical
Remote setup stipend
Internet Reimbursement
Gym Reimbursement

Job summary

Vultr seeks an AI Cluster Architect to design and refine large-scale GPU clusters within fixed power envelopes. You will optimize GPU density while accounting for compute, storage, networking, cooling, and facility constraints, balancing GPU power, fabric performance, and service density.

You will evaluate InfiniBand, RoCE, SpectrumX designs, model power usage, and develop templates for scalable deployments across sites, collaborating with vendors to enable 100k+ GPU deployments.

Qualifications

  • Experience designing large-scale GPU clusters within power constraints.
  • Familiarity with InfiniBand, RoCE, SpectrumX architectures.

Responsibilities

  • Architect large-scale GPU clusters within fixed site power budgets.
  • Model and validate power consumption across the cluster BOM.
  • Evaluate fabric networking architectures (InfiniBand, RoCE, SpectrumX) and multi-plane/topology options.
  • Determine network scale limits based on switch radix, link speed, topology, and blocking requirements.
  • Gather SKU-level power and thermal specs for GPUs, NICs, switches, DPUs, storage, and servers.
  • Develop power-aware cluster templates and capacity-planning models for multiple sites.
  • Document architecture decisions and lifecycle considerations for deployment.
  • Provide guidance on future-proofing with next-gen GPUs, NICs, or fabrics.
  • Collaborate with vendors on novel fabric architectures for 100k+ GPUs

Skills

GPU cluster design
Power-aware design
High-performance computing
Networking fabrics

Job description

Vultr seeks an AI Cluster Architect to design and refine large-scale GPU clusters within fixed power envelopes. You will optimize GPU density while accounting for compute, storage, networking, cooling, and facility constraints, balancing GPU power, fabric performance, and service density.

You will evaluate InfiniBand, RoCE, SpectrumX designs, model power usage, and develop templates for scalable deployments across sites, collaborating with vendors to enable 100k+ GPU deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Power-Aware GPU Cluster Architect for HPC
Power-Aware GPU Cluster Architect for HPC

Vultr • United States

On-site
USD 165,000 - 185,000
Excellent Medical Benefits
401(k) matching
Professional Development Reimbursement
+4
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
AI Cluster Architect
AI Cluster Architect

Vultr • United States

On-site
USD 165,000 - 185,000
Excellent Medical Benefits
401(k) matching
Professional Development Reimbursement
+4
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Cluster Architect
AI Cluster Architect

Webhosting • Northern (KY)

Hybrid
USD 165,000 - 185,000
Medical Benefits
401(k) Match
PD Reimbursement
+5
Remote GPU Cluster Architect for Scalable AI Infrastructure
Remote GPU Cluster Architect for Scalable AI Infrastructure

Nebius • United States

On-site
USD 184,000 - 318,000
Health insurance: 100% company-paid coverage
401(k) plan: Up to 4% company match
20 weeks paid parental leave for primary caregivers
+5
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI/HPC Cluster Architect - Scalable Data Center Design
AI/HPC Cluster Architect - Scalable Data Center Design

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits
Equal opportunity employer
Visa sponsorship not available