Lead AI Infrastructure & Distributed Systems Engineer

LinuxRecruit

Greater London

On-site

GBP 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LinuxRecruit is building a foundational AI system for physical intelligence, aiming to train large models efficiently on GPU clusters from a centrally located London base. The role demands hands-on coding, system design, and full ownership of the training infrastructure in an early-stage startup setting.

You will work at the intersection of HPC, GPU orchestration, and MLOps, crafting scalable systems, optimizing throughput, and driving architectural decisions that unlock rapid experimentation

Qualifications

  • Strong production background with AWS, Kubernetes, Slurm, PyTorch.
  • Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking.
  • Experience in MLOps, AI infrastructure, or HPC; robotics experience not required, though CV pipelines are advantageous.

Responsibilities

  • Own architectural decisions for core model training infrastructure from day one.
  • Design scalable distributed systems from scratch for large GPU compute clusters.
  • Code daily and reduce training bottlenecks to accelerate research cycles.

Skills

GPU compute optimisation
Cluster scheduling
High performance networking
MLOps
Distributed training

Tools

AWS
Kubernetes
Slurm
PyTorch
Distributed training frameworks

Job description

Building foundation model style AI for physical intelligence and pioneering a system that allows robots to learn complex physical tasks from just a single human demonstration (think "prompting" a robot in the physical world).

Having recently closed a $40M Seed round backed by premier global investors, (marking one of the largest robotics seed investments in European history), they are intentionally keeping their team small, elite, and talent dense. Their primary bottleneck is no longer collecting data, it is the raw speed, scale, and efficiency of training their models.

This business is looking for an elite engineer to "run the show" for their core model training infrastructure. Sitting at the intersection of high performance computing, GPU orchestration, and MLOps, this individual will professionalise the company's infrastructure, eliminate training bottlenecks, and drastically accelerate research cycles.

This is a deeply technical, hands on position. Rather than managing people or executing traditional DevOps/SRE maintenance, the successful candidate will be coding daily, designing scalable distributed systems from scratch, and squeezing maximum throughput out of large scale GPU compute clusters.

The founders are seeking a high agency engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation.

  • Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands‑on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential.
  • Relevant Background: Experience in MLOps, AI infrastructure, or HPC. Prior robotics experience is not required, though experience with computer vision pipelines is highly advantageous compared to text‑only LLM backgrounds.
  • Mindset: A desire to join a founding level team in person in central London, take full ownership of architectural decisions from day one, and capture substantial upside through equity.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI infrastructure engineer
AI infrastructure engineer

LinuxRecruit • Greater London

On-site
GBP 90,000 - 120,000
Competitive salary
Equity in early-stage startup
London lab
Founding AI Infrastructure Engineer
Founding AI Infrastructure Engineer

LinuxRecruit • Greater London

On-site
GBP 90,000 - 140,000
Senior AI Infra & DS Engineer — London
Senior AI Infra & DS Engineer — London

LinuxRecruit • Greater London

On-site
GBP 100,000 - 140,000
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Applied AI Engineer London
Applied AI Engineer London

Model ML • Greater London

On-site
GBP 110,000 - 150,000
Competitive salary + equity
Senior ML Engineer _TT
Senior ML Engineer _TT

PulseRise Technologies • Greater London

On-site
GBP 120,000 - 180,000
Research Engineer, Machine Learning (AI for Science)
Research Engineer, Machine Learning (AI for Science)

Generative • Greater London

On-site
GBP 90,000 - 150,000
Generative AI Engineer
Generative AI Engineer

Harrington Starr • Greater London

On-site
GBP 120,000 - 180,000
Founding Engineer - AI Infrastructure
Founding Engineer - AI Infrastructure

LinuxRecruit • Greater London

On-site
GBP 90,000 - 130,000
Principal AI Engineer
Principal AI Engineer

Harrington Starr • Greater London

On-site
GBP 70,000 - 100,000