HPC Cluster Architect

NexGen Cloud

United Kingdom

Hybrid

GBP 90,000 - 140,000

Full time

39 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive salary
25 days holiday
Remote or hybrid
Ownership & autonomy
Career growth

Job summary

NexGen Cloud is seeking a senior HPC architect to own the full architecture cycle for large GPU deployments, from first customer conversation to production handover. You’ll translate requirements into production-ready designs across compute, networking, storage, power and cooling.

You’ll act as a technical authority, owning end-to-end designs, guiding vendors, validating configurations, and supporting deployment and performance testing.

Qualifications

  • Proven experience in HPC or AI software stack design and delivery at scale.
  • Deep understanding of GPU software environments: CUDA, cuDNN, NCCL, driver stacks.

Responsibilities

  • Own end-to-end cluster architecture for large NVIDIA GPU deployments—from customer requirements through rack layouts, BOM, power & cooling design, to production handover.
  • Design high-performance network fabrics across compute, storage, and WAN—defining topology, oversubscription models, and scaling strategies.
  • Engage directly with OEMs and vendors—validating hardware configurations, reviewing quotes, and ensuring designs are technically sound and commercially optimised.
  • Provide technical oversight during deployment and bring‑up—supporting hardware validation, performance testing, and acting as escalation point for complex integration issues.
  • Act as a senior technical leader across Solutions Architecture, Cloud Engineering, and data centre partners—contributing to standardised reference designs and building out the HPC engineering function.

Skills

HPC Experience
GPU software stacks
Performance tuning
Containerisation & orchestration
OEM/vendor engagement
Documentation & diagrams
Customer-facing technical authority

Tools

SLURM
PBS
NVIDIA GPU Operator
Docker
Kubernetes

Job description

NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature.

We’re a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like.

This role exists because NexGen Cloud is winning large-scale dedicated GPU cluster contracts and needs someone who can own the full architecture cycle — from first customer conversation to production deployment. You’ll have direct ownership over cluster architecture across compute, networking, storage, and physical design — translating customer requirements into production-ready, commercially optimised GPU deployments.

This is a senior hands‑on role for someone who has lived and breathed HPC cluster design and wants to be the technical authority, not one voice in a committee. You’ll own designs end‑to‑end and see them go live.

WHAT YOU'LL BE DOING:

Rather than a long checklist, here’s what success in this role looks like:

  • Own end‑to‑end cluster architecture for large‑scale NVIDIA GPU deployments — from customer requirement through rack layouts, BOM, power and cooling design, to production handover
  • Design high‑performance network fabrics across compute (InfiniBand, RDMA, NVLink/NVSwitch), storage, and WAN — defining topology, oversubscription models, and scaling strategies
  • Engage directly with OEMs and vendors — validating hardware configurations, reviewing quotes, and ensuring designs are both technically sound and commercially optimised
  • Provide technical oversight during deployment and bring‑up — supporting hardware validation, performance testing, and acting as escalation point for complex integration issues
  • Act as a senior technical leader across Solutions Architecture, Cloud Engineering, and data centre partners — contributing to standardised reference designs and building out the HPC engineering function
ABOUT YOU:

We’re more interested in how you think and work than in a perfect CV. You’ll likely bring a combination of the following:

  • Proven experience in HPC or AI software stack design and delivery at scale — including workload profiling, scheduler configuration (SLURM, PBS, or equivalent), MPI/NCCL tuning, and distributed training frameworks such as PyTorch, JAX, or DeepSpeed.
  • Deep understanding of GPU software environments: CUDA, cuDNN, NCCL, driver stacks, and the tooling required to run large‑scale AI training and inference workloads reliably in production.
  • Hands‑on experience optimising AI and HPC workloads across multi‑GPU and multi‑node configurations — including profiling, bottleneck identification, and performance tuning at both the application and infrastructure layer.
  • Strong working knowledge of containerisation and orchestration in HPC/AI contexts: Docker, Kubernetes, NVIDIA GPU Operator, and container‑native workload management.
  • Background in an OEM, hyperscaler, neo‑cloud, or enterprise/research HPC environment, with demonstrable exposure to the full design‑to‑deployment lifecycle for GPU‑accelerated workloads.
  • Ability to produce clear, professional technical documentation and architecture diagrams suitable for both engineering and board‑level audiences.
  • Confident engaging with customers, vendors, and internal engineering teams as a technical authority — able to translate complex software and performance trade‑offs into clear, actionable decisions.

Nice to Have

  • Experience with large‑scale cluster performance benchmarking — NCCL tests, MLPerf, or equivalent — and familiarity with what good looks like across different GPU generations and topologies.
  • Exposure to MLOps tooling and AI platform layers: experiment tracking (MLflow, W&B), model serving frameworks (Triton, vLLM), and pipeline orchestration (Kubeflow, Airflow).
  • Familiarity with InfiniBand and high‑performance networking as it relates to distributed training performance — sufficient to engage credibly with network and infrastructure teams on topology and tuning decisions.
WHAT WE OFFER:
  • Competitive salary and annual discretionary bonus scheme
  • 25 days of holiday, plus public holidays
  • Flexible working arrangements (remote or hybrid, depending on role and location)
  • Real ownership and autonomy, with the trust to take initiative and experiment
  • The opportunity to make a visible, meaningful impact as we scale
  • Clear career progression and growth opportunities in a fast‑growing company
  • A collaborative, international culture built on trust, transparency, and ownership
  • The chance to help shape NexGen Cloud's team, culture, and future alongside ambitious, mission‑driven colleagues
MORE INFORMATION

Head over to our NexGen Cloud careers page to view current openings and follow us on LinkedIn and X to learn more about our journey, newest releases and hear exciting news in the neocloud space.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Project Manager
Technical Project Manager

NexGen Cloud • Greater London

On-site
GBP 75,000 - 110,000
Hybrid work
25 days holiday
Discretionary bonus
+2
Senior Infrastructure Engineer
Senior Infrastructure Engineer

NexGen Cloud • Greater London

Hybrid
GBP 90,000 - 120,000
Competitive salary
Annual discretionary bonus
25 days holiday
+2
Infrastructure Operations Engineer
Infrastructure Operations Engineer

NexGen Cloud • Greater London

Hybrid
GBP 75,000 - 110,000
Discretionary bonus
Flexible working
Wellbeing benefits
+1
Infrastructure Engineer - Finland
Infrastructure Engineer - Finland

NexGen Cloud • Greater London

Hybrid
GBP 65,000 - 95,000
Competitive salary and discretionary 1
Bonus scheme
25 days holiday
+3
Strategic Sales Lead, AI Natives
Strategic Sales Lead, AI Natives

NexGen Cloud Ltd • Greater London

Hybrid
GBP 65,000 - 90,000
Competitive salary
Annual discretionary bonus
Flexible working arrangements
+1
Business Development Manager
Business Development Manager

NexGen Cloud Ltd • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive salary and annual discretionary bonus
25 days of holiday plus public holidays
Flexible working arrangements
+1
Senior Content Manager
Senior Content Manager

NexGen Cloud • Greater London

On-site
GBP 65,000 - 90,000
Competitive salary
Private Medical Insurance
Flexible working (remote or hybrid)
+1
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure

NVIDIA • West of England

On-site
GBP 90,000 - 130,000
Senior Business Development Manager - Vertical Markets
Senior Business Development Manager - Vertical Markets

NexGen Cloud • Greater London

Hybrid
GBP 110,000 - 170,000
Equity participation
Private Medical Insurance
25 days holiday
Supply Chain Manager
Supply Chain Manager

NexGen Cloud • United Kingdom

Hybrid
GBP 70,000 - 110,000
Private Medical Insurance
Enhanced Pension Scheme
Flexible working arrangements
+2