Senior Solution Engineer – GPU and AI Infrastructure (CAF673F)

Referment

Greater London

On-site

GBP 90,000 - 150,000

Full time

26 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Referment is partnering with a cloud infrastructure provider to shape large-scale GPU platforms for UK enterprise customers. The role focuses on architecting and delivering scalable GPU clusters, with emphasis on performance, efficiency and cost trade-offs.

You will lead architecture workshops, design high-level and low-level documents, and translate workloads into bare-metal, Slurm or Kubernetes-based solutions across NVIDIA GPU platforms, NVLink/NVSwitch, InfiniBand fabrics and cooling

Qualifications

  • 5+ years in solution architecture, systems engineering or technical pre-sales focused on HPC, AI infrastructure or high-performance cloud platforms.
  • Deep knowledge of NVIDIA HGX or DGX systems, GPU interconnects and modern rack-scale GPU architectures.
  • Strong experience designing InfiniBand and/or RoCE networks, including congestion management and GPU-direct technologies.
  • Practical understanding of GPU workload orchestration through Kubernetes and associated NVIDIA operators, or bare-metal environments using Slurm, Ansible and Terraform.
  • A track record of producing robust HLDs, LLDs, network diagrams and itemised infrastructure designs.

Responsibilities

  • Design enterprise GPU clusters and produce clear high-level and low-level designs, rack and network diagrams, and detailed bills of materials.
  • Shape scale-up and scale-out architectures across NVIDIA GPU platforms, NVLink and NVSwitch, high-speed InfiniBand and RoCE fabrics, storage, power and cooling.
  • Translate customer workload, performance and commercial requirements into practical bare-metal, Slurm or Kubernetes-based solutions.
  • Lead technical discovery sessions, architecture workshops and executive presentations, and contribute to complex proposals and RFP responses.
  • Plan and oversee proofs of concept, using appropriate benchmarking to validate throughput, latency and distributed training performance.
  • Work closely with enterprise customers, hardware vendors and internal engineering teams, feeding technical insight back into the platform roadmap.

Skills

HPC architecture
GPGPU platforms
Kubernetes orchestration
InfiniBand networks
Pre-sales experience

Tools

Kubernetes
Slurm
Ansible
Terraform
NVIDIA HGX/DGX

Job description

Referment is working with a cloud infrastructure provider that supports enterprise AI, high-performance computing and cloud-native workloads. The business is looking for a senior technical specialist to shape large-scale GPU platforms for customers across the UK.

The Role
  • Design enterprise GPU clusters and produce clear high-level and low-level designs, rack and network diagrams, and detailed bills of materials.
  • Shape scale-up and scale-out architectures across NVIDIA GPU platforms, NVLink and NVSwitch, high-speed InfiniBand and RoCE fabrics, storage, power and cooling.
  • Translate customer workload, performance and commercial requirements into practical bare-metal, Slurm or Kubernetes-based solutions.
  • Lead technical discovery sessions, architecture workshops and executive presentations, and contribute to complex proposals and RFP responses.
  • Plan and oversee proofs of concept, using appropriate benchmarking to validate throughput, latency and distributed training performance.
  • Work closely with enterprise customers, hardware vendors and internal engineering teams, feeding technical insight back into the platform roadmap.
What We’re Looking For
  • At least five years in solution architecture, systems engineering or technical pre-sales focused on HPC, AI infrastructure or high-performance cloud platforms.
  • Deep knowledge of NVIDIA HGX or DGX systems, GPU interconnects and modern rack-scale GPU architectures.
  • Strong experience designing InfiniBand and/or RoCE networks, including congestion management and GPU-direct technologies.
  • Practical understanding of GPU workload orchestration through Kubernetes and associated NVIDIA operators, or bare-metal environments using technologies such as Slurm, Ansible and Terraform.
  • A track record of producing robust HLDs, LLDs, network diagrams and itemised infrastructure designs.
  • Confident communication with CTOs, infrastructure leaders and ML engineers, with the judgement to explain hardware, networking and cost trade-offs.
  • Awareness of high-density data-centre power, cooling and storage considerations.
Relevant Desirable Experience

NVIDIA AI infrastructure, networking or InfiniBand certifications would be useful, as would experience with performance testing for distributed AI workloads. A relevant degree is welcome, although equivalent practical experience is equally valuable.

This could suit a Senior Solutions Architect, HPC Systems Engineer or technical pre-sales specialist who has designed GPU clusters and wants to remain close to both customers and engineering. You must be based in the UK.

#Referment

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Solution Engineer – GPU and AI Infrastructure (F673F3F)
Senior Solution Engineer – GPU and AI Infrastructure (F673F3F)

Referment • Greater London

On-site
GBP 90,000 - 120,000
Senior GPU Solutions Architect – Enterprise HPC UK
Senior GPU Solutions Architect – Enterprise HPC UK

Referment • Greater London

On-site
GBP 90,000 - 150,000
Solutions Architect
Solutions Architect

WNTD • England

Hybrid
GBP 70,000 - 90,000
Hybrid working model
Cutting-edge technology exposure
Opportunity for professional growth
Technical Solutions Architect – Investors
Technical Solutions Architect – Investors

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
Senior GPU & AI Infra Architect — Remote, 4-Day Week
Senior GPU & AI Infra Architect — Remote, 4-Day Week

Civo Ltd • United Kingdom

Hybrid
GBP 110,000 - 170,000
4-day week
Uncapped holidays
Remote work environment
GPU Infrastructure Lead - Systems Integrator
GPU Infrastructure Lead - Systems Integrator

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 140,000 - 170,000
Full Benefits
Presales Engineer
Presales Engineer

Hamilton Barnes ? • Greater London

Hybrid
GBP 70,000 - 90,000
Senior GPU & AI Infrastructure Architect
Senior GPU & AI Infrastructure Architect

Referment • Greater London

On-site
GBP 90,000 - 120,000
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Xapply • United Kingdom

Remote
GBP 90,000 - 130,000
Senior Solutions Architect, Higher Education and Research
Senior Solutions Architect, Higher Education and Research

NVIDIA Corporation • Reading

Hybrid
GBP 85,000 - 135,000