Head of Physical Infrastructure

Callosum

Greater London

On-site

GBP 120,000 - 160,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive Salary
Equity & Ownership
Private healthcare
Visa sponsorship and relocation
London office

Job summary

Callosum is hiring to own its physical compute infrastructure across sites and vendors, from assessment to reliable operation of clusters. You will scope sites, design cluster infrastructure, and lead commissioning of new hardware while building a team to keep infrastructure healthy and available.

As the footprint grows, you will establish operational systems for telemetry, monitoring, capacity planning, and lifecycle management, driving reliability and observability in a rapidly evolving AI

Qualifications

  • Experience designing, deploying, or operating high-density compute infrastructure, HPC clusters, AI infrastructure, or similarly complex physical systems.
  • Strong understanding of how power, cooling, rack design, networking, storage, and compute interact to determine cluster capabilities.
  • Demonstrated ability to drive complex infrastructure deployments across technical teams, facilities, hardware vendors and other external stakeholders.

Responsibilities

  • Scope new cluster deployments across site selection, power, cooling, rack layout, networking, storage and capacity requirements.
  • Coordinate delivery and commissioning across facilities, utilities, OEMs, networking and storage vendors.
  • Build the operational infrastructure for the fleet, including hardware telemetry, health monitoring and incident response.
  • Build and lead the physical infrastructure team responsible for deploying new clusters and maintaining reliability.

Skills

High-density compute
HPC clusters
AI infrastructure
Infrastructure leadership

Job description

About Us

We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity. Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next. The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve. Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost. Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor. In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence. We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.

About the Role

Callosum believes that the next generation of AI infrastructure will be built from a much more diverse set of hardware than the clusters of today. Scaling that infrastructure creates a new physical systems problem: different accelerators bring different power densities, cooling requirements, rack constraints, network topologies, vendor dependencies, and deployment models. We want to find ways to turn those increasingly complex inputs into physical infrastructure that can be deployed repeatedly and scaled. This role owns Callosum’s physical compute infrastructure, from the first assessment of a potential deployment through to the reliable operation of the resulting cluster. You will scope sites, coordinate the design cluster infrastructure and deployment across vendors and facilities, and lead commissioning of new hardware. As our footprint grows, you will build and lead the team responsible for keeping that infrastructure healthy, observable, and available, establishing the operational systems that turn individual deployments into a reliable fleet.

What You’ll Do
  • Scope new cluster deployments across site selection, power, cooling, rack layout, networking, storage, fibre connectivity, and capacity requirements
  • Coordinate delivery and commissioning across facilities, utilities, OEMs, networking and storage vendors, ensuring the physical and technical dependencies come together
  • Build the operational infrastructure for the fleet, including hardware telemetry, health monitoring, incident response, capacity planning, and hardware lifecycle management
  • Build and lead the physical infrastructure team responsible for deploying new clusters and maintaining the reliability and availability of the resources we operate
What Sets You Apart
  • Experience designing, deploying, or operating high-density compute infrastructure, HPC clusters, AI infrastructure, or similarly complex physical systems
  • Strong understanding of how power, cooling, rack design, networking, storage, and compute interact to determine the capabilities and constraints of a cluster
  • Demonstrated ability to drive complex infrastructure deployments across technical teams, facilities, hardware vendors, network providers, and other external stakeholders
  • Experience building and leading infrastructure teams, with strong operational instincts around reliability, observability, capacity, maintenance, and failure management
What We Offer
  • Competitive Salary, determined by skills and experience
  • Equity & Ownership
  • Private healthcare
  • We offer Visa sponsorship and relocation benefits to hire the best in the world
  • We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us

We’re committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Physical Infrastructure
Head of Physical Infrastructure

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 130,000
Competitive salary
Equity
Private healthcare
+2
Head of Physical Infrastructure
Head of Physical Infrastructure

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 90,000 - 120,000
Visa sponsorship
Relocation benefits
Private healthcare
+2
Chief of Staff, CTO
Chief of Staff, CTO

Callosum • Greater London

On-site
GBP 110,000 - 150,000
Competitive Salary
Equity & Ownership
Private healthcare
+2
Networking & Interconnect Systems - Member of Technical Staff
Networking & Interconnect Systems - Member of Technical Staff

Callosum • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity & Ownership
Private healthcare
+3
Security Lead - Member of Technical Staff
Security Lead - Member of Technical Staff

Callosum • Greater London

On-site
GBP 120,000 - 180,000
Equity
Private healthcare
Visa sponsorship
+1
Compliance Lead - Member of Technical Staff
Compliance Lead - Member of Technical Staff

Callosum • Greater London

On-site
GBP 90,000 - 140,000
Equity
Private healthcare
Visa sponsorship
+2
Inference System & Performance - Member of Technical Staff
Inference System & Performance - Member of Technical Staff

Callosum • Greater London

On-site
GBP 101,000 - 192,000
Competitive salary
Equity & Ownership
Private healthcare
+2
Site Reliability - Member of Technical Staff
Site Reliability - Member of Technical Staff

Callosum • Greater London

On-site
GBP 90,000 - 150,000
Equity & Ownership
Private healthcare
Visa sponsorship & relocation
+1
Networking & Interconnect Systems - Member of Technical Staff
Networking & Interconnect Systems - Member of Technical Staff

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 150,000
Equity
Private healthcare
Visa sponsorship
+1
ML Research Scientist - Member of Technical Staff
ML Research Scientist - Member of Technical Staff

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive salary
Equity & Ownership
Private healthcare
+2