AI Cluster Architect

Webhosting

Northern (KY)

On-site

USD 165,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical Benefits
401(k) Match
PD Reimbursement
PTO & Holidays
Sabbatical
Remote setup stipend
Internet Reimbursement
Gym Reimbursement

Job summary

Vultr seeks an AI Cluster Architect to design and refine large-scale GPU clusters within fixed power envelopes. You will optimize GPU density while accounting for compute, storage, networking, cooling, and facility constraints, balancing GPU power, fabric performance, and service density.

You will evaluate InfiniBand, RoCE, SpectrumX designs, model power usage, and develop templates for scalable deployments across sites, collaborating with vendors to enable 100k+ GPU deployments.

Qualifications

  • Experience designing large-scale GPU clusters within power constraints.
  • Familiarity with InfiniBand, RoCE, SpectrumX architectures.

Responsibilities

  • Architect large-scale GPU clusters within fixed site power budgets.
  • Model and validate power consumption across the cluster BOM.
  • Evaluate fabric networking architectures (InfiniBand, RoCE, SpectrumX) and multi-plane/topology options.
  • Determine network scale limits based on switch radix, link speed, topology, and blocking requirements.
  • Gather SKU-level power and thermal specs for GPUs, NICs, switches, DPUs, storage, and servers.
  • Develop power-aware cluster templates and capacity-planning models for multiple sites.
  • Document architecture decisions and lifecycle considerations for deployment.
  • Provide guidance on future-proofing with next-gen GPUs, NICs, or fabrics.
  • Collaborate with vendors on novel fabric architectures for 100k+ GPUs

Skills

GPU cluster design
Power-aware design
High-performance computing
Networking fabrics

Job description

Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation. Founded by David Aninowsky and self-funded for over a decade, Vultr has grown to become the world’s largest privately-held cloud infrastructure company.

Vultr Cares
  • Excellent Medical Benefits w/ 100% company-paid premiums for employee only plan + 100% company-paid dental & vision premiums
  • 401(k) plan that matches 100% up to 4% with immediate vesting
  • Professional Development Reimbursement of $2,500 each year
  • 11 Holidays + Paid Time Off Accrual + Rollover Plan + take your birthday off
  • Commitment matters to Vultr! Increased PTO at 3 year & 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year
  • $500 first year remote office setup + $400 each following year for new equipment
  • Internet reimbursement up to $75 per month
  • Gym membership reimbursement up to $50 per month
Join Vultr

Vultr is looking for an AI Cluster Architect who will be responsible for creating and refining large-scale GPU cluster architectures within strict power and infrastructure limits. This role focuses heavily on power-aware design: starting from a fixed power envelope, the architect determines the optimal number of GPUs while accounting for the full stack of services needing to be deployed-compute nodes, storage systems, networking fabric, cooling, and facility constraints. This role requires deep experience navigating heterogeneous environments, multiple generations of hardware, and end user requirements.

The architect must understand how different GPU SKUs, NICs, switches, and fabrics interact at scale, including their individual and aggregate power and thermal characteristics. They will evaluate multi-plane, rail-optimized, and tiered fabric designs across technologies like InfiniBand, RoCE, and SpectrumX to ensure the networking architecture supports the intended GPU count without overrunning facility limits or switch radix and/or topology constraints. This role balances customer-specific requirements for compute, storage, and service density, ensuring that the final cluster design maintains acceptable levels of GPU and fabric performance, while maximizing the number of usable GPUs within the total power budget.

Key Responsibilities

Architect large-scale GPU clusters within fixed site power budgets that optimizes for maximum GPU density while reserving necessary headroom for compute services, storage, and networking.

Model and validate power consumption across the full cluster bill of materials (GPUs, CPUs, NICs, switches, fabric components, storage, and facility limits).

Evaluate tradeoffs across multiple fabric networking architectures (InfiniBand, RoCE, SpectrumX) as well as multi-plane, 2-tier/3-tier, and rail-optimized topologies.

Determine network scale limits based on switch radix, link speed, topology, and blocking requirements.

Gather, interpret, and maintain detailed SKU-level power and thermal specifications for GPUs, NICs, switches, DPUs, storage, and server platforms.

Develop power-aware cluster configuration templates and capacity-planning models that can scale across sites with varying constraints and allow for quick iteration and ideation.

Document architecture, design choices, tradeoff analyses, and operational considerations for deployment and lifecycle management.

Provide guidance on future-proofing, including the ability to incorporate next-gen GPUs, NICs, or fabrics.

Collaborate with vendors on novel fabric architectures that enable large-scale cluster deployments (100k+ GPUs)

Compensation

$165,000 – $185,000

This salary can vary based on location, years of experience, background and skill set.

Inclusion & Privacy

We are an equal opportunity employer and are committed to creating an inclusive environment for all employees. We welcome applications from individuals of all backgrounds and experiences, and we prohibit discrimination based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected status under applicable laws. Vultr will consider qualified applicants with arrest or conviction records in accordance with applicable laws and will not conduct a background check until after an offer of employment has been extended and accepted.

We also take your privacy seriously. We handle personal information responsibly and follow applicable laws, including U.S. privacy rules and India's Digital Personal Data Protection Act, 2023. Your data is used only for legitimate business purposes and is protected with proper security measures.

Where allowed by law, applicants may request details about the data we collect, access or delete their information, withdraw consent for its use, and opt out of nonessential communications. For more details, please see our Privacy Policy .

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Cluster Architect
AI Cluster Architect

Vultr • United States

On-site
USD 165,000 - 185,000
Excellent Medical Benefits
401(k) matching
Professional Development Reimbursement
+4
Staff AI/ML Infrastructure Engineer
Staff AI/ML Infrastructure Engineer

Vultr • United States

Remote
USD 145,000 - 160,000
100% company-paid insurance premiums
401(k) plan with matching
Professional Development Reimbursement
+4
Strategic Technical Account Manager GPU
Strategic Technical Account Manager GPU

Vultr • United States

Remote
USD 115,000 - 140,000
100% company-paid insurance premiums
401(k) match up to 4%
Professional Development Reimbursement
+4
Data Center Site Selection Manager
Data Center Site Selection Manager

Vultr • United States

Remote
USD 110,000 - 130,000
100% company-paid insurance premiums
401(k) plan with matching
Professional Development Reimbursement
+4
Senior Technical Project Manager, Data Center & Network Delivery
Senior Technical Project Manager, Data Center & Network Delivery

Vultr • United States

On-site
USD 110,000 - 140,000
100% company-paid insurance premiums for medical, dental, and vision plans
401(k) plan with 100% match up to 4%
Professional Development Reimbursement of $2,500 each year
+6
Principal Technical Product Manager, Strategic Accounts
Principal Technical Product Manager, Strategic Accounts

Socket.dev • United States

Hybrid
USD 200,000 - 230,000
Data Center Technican (Eagan, MN)
Data Center Technican (Eagan, MN)

Vultr • Eagan (MN)

On-site
USD 60,000 - 95,000
100% company-paid insurance premiums
401(k) plan with company match
Professional development reimbursement
+2
Infrastructure Production Engineer
Infrastructure Production Engineer

Vultr • United States

On-site
USD 60,000 - 80,000
Premium health insurance
401(k) with company match
Professional development reimbursement
+6
Data Center Technician (Kansas City)
Data Center Technician (Kansas City)

Vultr • Kansas City (MO)

On-site
USD 60,000 - 75,000
100% company-paid insurance premiums
401(k) matching
Professional Development Reimbursement
+5
Infrastructure Production Engineer
Infrastructure Production Engineer

Webhosting • Northern (KY)

Hybrid
USD 60,000 - 80,000
Insurance
401(k) matching
Professional development reimbursement
+5