GPU Infrastructure Lead - Systems Integrator

Hamilton Barnes Associates Limited

Greater London

In loco

GBP 140.000 - 170.000

Tempo pieno

14 giorni+
Generatore di candidature

Distinguiti per questa posizione — genera un curriculum e una lettera di presentazione personalizzati in circa un minuto.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Full Benefits

Descrizione del lavoro

Hamilton Barnes Associates Limited is seeking an experienced infrastructure leader to own the GPU compute infrastructure function. You will design validation and certification frameworks, and build the tooling to scale testing across large GPU clusters.

This role blends hands-on engineering with technical leadership, shaping standards, collaborating with founders, and growing a team to deliver reliable, high-performance compute capacity for AI workloads.

Competenze

  • Experience deploying GPU clusters at varying scales, from smaller node deployments to environments with hundreds or thousands of GPUs, across both air-cooled and liquid-cooled infrastructure.
  • Strong experience troubleshooting complex infrastructure issues, including cabling and topology errors, firmware mismatches, failing optics, thermal throttling, and production performance challenges.
  • Experience automating infrastructure-as-code deployments, monitoring and alerting stacks, and hardware acceptance and regression testing.
  • Deep familiarity with the NVIDIA technology stack, including HGX platforms, NVLink/NVSwitch, CUDA-level debugging, and NCCL performance tuning.
  • Strong leadership skills with the ability to set technical direction, solve complex problems, and guide engineering teams.
  • Excellent communication skills with the ability to engage effectively with non-infrastructure stakeholders, including traders, lawyers, capacity providers, and regulators.

Mansioni

  • Design and run cluster validation and certification: performance benchmarks, interconnect testing (NCCL, InfiniBand/RoCE), thermal and power verification, availability monitoring against SLAs.
  • Build automation for cluster deployment, health checks, and continuous testing so certification scales without headcount scaling with it.
  • Set the technical standards for what deliverable compute means: node configs, network topologies, storage, cooling envelopes.
  • Work directly with capacity providers during onboarding, from site walkthroughs to acceptance testing.
  • Feed what you learn on the ground back into our contract specs, index methodology, and product roadmap.
  • Hire and lead the infrastructure engineering team as we grow.

Descrizione del lavoro

Are you looking for an exciting new opportunity?

Join a pioneering compute infrastructure technology company building the platforms and tools that power the rapidly evolving AI compute market. The organisation develops financial and settlement infrastructure for compute, including technology that connects buyers with GPU capacity, while building automated systems to validate, benchmark, and certify large-scale GPU clusters.

The successful candidate will take end-to-end ownership of the GPU infrastructure function, developing the frameworks and automation used to validate, benchmark, and certify large-scale GPU clusters. Combining hands-on engineering with technical leadership, they will build deployment and testing tooling, establish infrastructure standards, shape technical strategy alongside the founders, and ultimately build and lead the team responsible for the function.

Responsibilities:
  • Design and run cluster validation and certification: performance benchmarks, interconnect testing (NCCL, InfiniBand/RoCE), thermal and power verification, availability monitoring against SLAs.
  • Build automation for cluster deployment, health checks, and continuous testing so certification scales without headcount scaling with it.
  • Set the technical standards for what "deliverable compute" means: node configs, network topologies, storage, cooling envelopes.
  • Work directly with capacity providers (neoclouds, data centre operators) during onboarding, from site walkthroughs to acceptance testing.
  • Feed what you learn on the ground back into our contract specs, index methodology, and product roadmap.
  • Hire and lead the infrastructure engineering team as we grow.
Skills/Must have:
  • Experience deploying GPU clusters at varying scales, from smaller node deployments to environments with hundreds or thousands of GPUs, across both air-cooled and liquid-cooled infrastructure. DLC experience is highly advantageous.
  • Strong experience troubleshooting complex infrastructure issues, including cabling and topology errors, firmware mismatches, failing optics, thermal throttling, and production performance challenges.
  • Experience automating infrastructure-as-code deployments, monitoring and alerting stacks, and hardware acceptance and regression testing.
  • Deep familiarity with the NVIDIA technology stack, including HGX platforms, NVLink/NVSwitch, CUDA-level debugging, and NCCL performance tuning.
  • Strong leadership skills with the ability to set technical direction, solve complex problems, and guide engineering teams.
  • Excellent communication skills with the ability to engage effectively with non-infrastructure stakeholders, including traders, lawyers, capacity providers, and regulators.
Benefits:
  • Full Benefits
Salary:
  • £150,000 Base Salary
Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Technical Solutions Architect – Investors
Technical Solutions Architect – Investors

Hamilton Barnes Associates Limited • Greater London

In loco
GBP 110.000 - 150.000
Network Architect (GPU AI Data Centres)
Network Architect (GPU AI Data Centres)

Experis UK • Belfast City District

Ibrido
GBP 90.000 - 130.000
Hybrid work model
On-site facilities
Solutions Architect
Solutions Architect

WNTD • England

In loco
GBP 70.000 - 90.000
Hybrid working model
Cutting-edge technology exposure
Opportunity for professional growth
Network Engineer
Network Engineer

asobbi • Regno Unito

In loco
GBP 53.195 - 68.394
Highly competitive package with equity
Dynamic progression plan
Human-first flexibility
GPU Infrastructure Lead: Scale, Certification, and Automation
GPU Infrastructure Lead: Scale, Certification, and Automation

Hamilton Barnes Associates Limited • Greater London

In loco
GBP 140.000 - 170.000
Full Benefits
Founding GPU Engineer
Founding GPU Engineer

Fuse Energy • Greater London

In loco
GBP 90.000 - 130.000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
HPC Infrastructure Site Reliability Engineer
HPC Infrastructure Site Reliability Engineer

Radiant • Greater London

In loco
GBP 90.000 - 140.000
Lead GPU Infrastructure Architect for Scalable AI Clusters
Lead GPU Infrastructure Architect for Scalable AI Clusters

Hamilton Barnes Associates Limited • Greater London

In loco
GBP 110.000 - 150.000
NOC Engineer - AI infrastructure
NOC Engineer - AI infrastructure

Hamilton Barnes • West of England

In loco
GBP 30.000 - 50.000
Head of Data Centres (Development & Strategy)
Head of Data Centres (Development & Strategy)

Troi • England

In loco
GBP 180.000 - 280.000
Competitive Salary
Unlimited Holiday Policy
Advanced Pension Scheme
+3