GPU Infrastructure Lead: Scale, Certification, and Automation

Hamilton Barnes Associates Limited

Greater London

On-site

GBP 140,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Full Benefits

Job summary

Hamilton Barnes Associates Limited is seeking an experienced infrastructure leader to own the GPU compute infrastructure function. You will design validation and certification frameworks, and build the tooling to scale testing across large GPU clusters.

This role blends hands-on engineering with technical leadership, shaping standards, collaborating with founders, and growing a team to deliver reliable, high-performance compute capacity for AI workloads.

Qualifications

  • Experience deploying GPU clusters at varying scales, from smaller node deployments to environments with hundreds or thousands of GPUs, across both air-cooled and liquid-cooled infrastructure.
  • Strong experience troubleshooting complex infrastructure issues, including cabling and topology errors, firmware mismatches, failing optics, thermal throttling, and production performance challenges.
  • Experience automating infrastructure-as-code deployments, monitoring and alerting stacks, and hardware acceptance and regression testing.
  • Deep familiarity with the NVIDIA technology stack, including HGX platforms, NVLink/NVSwitch, CUDA-level debugging, and NCCL performance tuning.
  • Strong leadership skills with the ability to set technical direction, solve complex problems, and guide engineering teams.
  • Excellent communication skills with the ability to engage effectively with non-infrastructure stakeholders, including traders, lawyers, capacity providers, and regulators.

Responsibilities

  • Design and run cluster validation and certification: performance benchmarks, interconnect testing (NCCL, InfiniBand/RoCE), thermal and power verification, availability monitoring against SLAs.
  • Build automation for cluster deployment, health checks, and continuous testing so certification scales without headcount scaling with it.
  • Set the technical standards for what deliverable compute means: node configs, network topologies, storage, cooling envelopes.
  • Work directly with capacity providers during onboarding, from site walkthroughs to acceptance testing.
  • Feed what you learn on the ground back into our contract specs, index methodology, and product roadmap.
  • Hire and lead the infrastructure engineering team as we grow.

Job description

Hamilton Barnes Associates Limited is seeking an experienced infrastructure leader to own the GPU compute infrastructure function. You will design validation and certification frameworks, and build the tooling to scale testing across large GPU clusters.

This role blends hands-on engineering with technical leadership, shaping standards, collaborating with founders, and growing a team to deliver reliable, high-performance compute capacity for AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead GPU Infrastructure Architect for Scalable AI Clusters
Lead GPU Infrastructure Architect for Scalable AI Clusters

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
GPU Infrastructure Lead - Systems Integrator
GPU Infrastructure Lead - Systems Integrator

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 140,000 - 170,000
Full Benefits
AI Data Center Engineer: GPU HPC & Networking
AI Data Center Engineer: GPU HPC & Networking

Hamilton Barnes Associates Limited • United Kingdom

On-site
GBP 45,000 - 75,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Uncover • Greater London

Hybrid
GBP 120,000 - 190,000
HPC Linux Engineer - Scale Compute & Automation
HPC Linux Engineer - Scale Compute & Automation

Hamilton Barnes Associates Limited • United Kingdom

On-site
GBP 85,000 - 110,000
Engineering-first culture
Growth opportunities across Linux &amp
Growth opportunities across Linux &amp
GPU Infra Engineer — Scale, Automation & AI Compute
GPU Infra Engineer — Scale, Automation & AI Compute

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Remote Senior Pre-Sales Architect - GPU AI Infra
Remote Senior Pre-Sales Architect - GPU AI Infra

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 135,000 - 165,000
Tailored Bonus Structure
Remote working
Senior GPU & AI Infra Architect — Remote, 4-Day Week
Senior GPU & AI Infra Architect — Remote, 4-Day Week

Civo Ltd • United Kingdom

Hybrid
GBP 110,000 - 170,000
4-day week
Uncapped holidays
Remote work environment
Technical Solutions Architect – Investors
Technical Solutions Architect – Investors

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000