Senior Solution Engineer – GPU and AI Infrastructure (F673F3F)

Referment

Greater London

On-site

GBP 90,000 - 120,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Referment is partnering with a cloud infrastructure provider to shape large-scale GPU platforms for customers across the UK. The role focuses on building enterprise GPU clusters, defining robust architectures and detailed BOMs.

You will translate workloads into Kubernetes-based or bare-metal solutions, lead discovery sessions, and drive proofs of concept with benchmarking to validate performance and cost considerations.

Qualifications

  • Minimum five years in solution architecture, systems engineering or technical pre-sales for HPC/AI infra.
  • Deep knowledge of NVIDIA HGX or DGX systems and GPU interconnects.
  • Experience with InfiniBand/RoCE networks and GPU-direct technologies.
  • Hands-on with Kubernetes for workload orchestration or bare-metal with Slurm/Ansible/Terraform.
  • Proven ability to produce end-to-end designs and diagrams.

Responsibilities

  • Design enterprise GPU clusters with high-level & low-level designs, rack and network diagrams, and detailed bills of materials.
  • Shape scale-up and scale-out architectures across NVIDIA GPU platforms, NVLink and NVSwitch, high-speed InfiniBand and RoCE fabrics, storage, power and cooling.
  • Translate customer workload, performance and commercial requirements into practical bare-metal, Slurm or Kubernetes-based solutions.
  • Lead technical discovery sessions, architecture workshops and executive presentations, and contribute to complex proposals and RFP responses.
  • Plan and oversee proofs of concept, using appropriate benchmarking to validate throughput, latency and distributed training performance.
  • Work closely with enterprise customers, hardware vendors and internal engineering teams, feeding technical insight back into the platform roadmap.

Skills

Solution architecture
Systems engineering
Technical pre Sales
NVIDIA HGX/DGX knowledge
GPU interconnects
InfiniBand/RoCE networking
Kubernetes
Slurm
Ansible
Terraform
HLD/LLD design
Executive communication

Education

Degree in a relevant field

Tools

NVIDIA HGX/DGX systems
InfiniBand switches
RoCE fabrics

Job description

Referment is working with a cloud infrastructure provider that supports enterprise AI, high-performance computing and cloud-native workloads. The business is looking for a senior technical specialist to shape large-scale GPU platforms for customers across the UK.

The Role
  • Design enterprise GPU clusters and produce clear high-level and low-level designs, rack and network diagrams, and detailed bills of materials.
  • Shape scale-up and scale-out architectures across NVIDIA GPU platforms, NVLink and NVSwitch, high-speed InfiniBand and RoCE fabrics, storage, power and cooling.
  • Translate customer workload, performance and commercial requirements into practical bare-metal, Slurm or Kubernetes-based solutions.
  • Lead technical discovery sessions, architecture workshops and executive presentations, and contribute to complex proposals and RFP responses.
  • Plan and oversee proofs of concept, using appropriate benchmarking to validate throughput, latency and distributed training performance.
  • Work closely with enterprise customers, hardware vendors and internal engineering teams, feeding technical insight back into the platform roadmap.
What We’re Looking For
  • At least five years in solution architecture, systems engineering or technical pre-sales focused on HPC, AI infrastructure or high-performance cloud platforms.
  • Deep knowledge of NVIDIA HGX or DGX systems, GPU interconnects and modern rack-scale GPU architectures.
  • Strong experience designing InfiniBand and/or RoCE networks, including congestion management and GPU-direct technologies.
  • Practical understanding of GPU workload orchestration through Kubernetes and associated NVIDIA operators, or bare-metal environments using technologies such as Slurm, Ansible and Terraform.
  • A track record of producing robust HLDs, LLDs, network diagrams and itemised infrastructure designs.
  • Confident communication with CTOs, infrastructure leaders and ML engineers, with the judgement to explain hardware, networking and cost trade-offs.
  • Awareness of high-density data-centre power, cooling and storage considerations.
Relevant Desirable Experience

NVIDIA AI infrastructure, networking or InfiniBand certifications would be useful, as would experience with performance testing for distributed AI workloads. A relevant degree is welcome, although equivalent practical experience is equally valuable.

This could suit a Senior Solutions Architect, HPC Systems Engineer or technical pre-sales specialist who has designed GPU clusters and wants to remain close to both customers and engineering. You must be based in the UK.

#Referment

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Solutions Architect
Solutions Architect

WNTD • England

Hybrid
GBP 70,000 - 90,000
Hybrid working model
Cutting-edge technology exposure
Opportunity for professional growth
Technical Solutions Architect – Investors
Technical Solutions Architect – Investors

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
GPU Infrastructure Lead - Systems Integrator
GPU Infrastructure Lead - Systems Integrator

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 140,000 - 170,000
Full Benefits
Senior GPU & AI Infra Architect — Remote, 4-Day Week
Senior GPU & AI Infra Architect — Remote, 4-Day Week

Civo Ltd • United Kingdom

Hybrid
GBP 110,000 - 170,000
4-day week
Uncapped holidays
Remote work environment
HPC Cluster Architect
HPC Cluster Architect

NexGen Cloud • United Kingdom

Hybrid
GBP 90,000 - 140,000
Competitive salary
25 days holiday
Remote or hybrid
+2
Senior GPU & AI Infrastructure Architect
Senior GPU & AI Infrastructure Architect

Referment • Greater London

On-site
GBP 90,000 - 120,000
Presales Engineer
Presales Engineer

Hamilton Barnes ? • Greater London

Hybrid
GBP 70,000 - 90,000
Senior Solutions Architect, Higher Education and Research
Senior Solutions Architect, Higher Education and Research

NVIDIA Corporation • Reading

Hybrid
GBP 85,000 - 135,000
Senior Solutions Architect, Higher Education and Research - Open Models and LLM
Senior Solutions Architect, Higher Education and Research - Open Models and LLM

NVIDIA AI • Cambridge

On-site
GBP 110,000 - 170,000
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC

NVIDIA • United Kingdom

On-site
GBP 110,000 - 170,000