Founding AI Infrastructure Engineer

Imperial College London

Greater London

On-site

GBP 90,000 - 130,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Founding-team influence
Small team culture

Job summary

Imperial College London in the United Kingdom invites applications from an exceptionally skilled engineer to build and own our AI training infrastructure. You will work across research and production environments to design scalable systems that run on hundreds of GPUs across cloud providers, emphasizing performance, reliability, and maintainability.

You will collaborate with researchers and platform engineers to optimize throughput, memory management and networking while advancing MLOps tooling

Qualifications

  • Experience with large-scale distributed training and PyTorch.
  • Strong background in HPC, GPU clusters and memory management.
  • Proficient Python and production-grade software engineering.
  • Experience with MLOps tooling and CI/CD.
  • Ability to work across research and production environments.
  • Familiarity with cloud platforms (AWS) and GPU cloud providers.
  • Experience with computer vision or multimodal AI is a plus.

Responsibilities

  • Design and own large-scale AI training infrastructure.
  • Scale distributed model training across hundreds of GPUs and multiple clouds.
  • Optimize training throughput and GPU utilization.

Skills

PyTorch
Distributed training
Kubernetes
Slurm
AWS
HPC
MLOps tooling
Python
GPU orchestration

Tools

Kubernetes
Slurm
GPU orchestration
AWS

Job description

What we're looking for:

You'll likely come from an MLOps, AI infrastructure, platform engineering or HPC background, but titles matter less than experience.

  • Large-scale distributed training.
  • PyTorch and modern deep learning frameworks.
  • Kubernetes, Slurm or GPU orchestration platforms.
  • AWS and specialist GPU cloud providers.
  • High-performance computing and distributed systems.
  • Training optimisation, memory management and networking.
  • MLOps tooling, CI/CD and experimentation frameworks.
  • Python and production-grade software engineering.
  • Computer vision, multimodal AI or transformer architectures.
Bonus points if you've worked on:
  • Foundation models.
  • Research infrastructure.
  • Simulation systems.
  • Robotics or embodied AI.
Who you are:

You're deeply technical and happiest when solving difficult engineering problems. You thrive in small, ambitious teams. You enjoy building from scratch rather than maintaining legacy systems. You can move comfortably between research and production. You care about performance, elegance and impact. You want ownership, autonomy and the opportunity to shape something significant.

Why join us?
  • Join one of Europe's most exciting AI and robotics startups.
  • Work alongside leading researchers from Imperial College London.
  • Meaningful equity and genuine founding-team influence.
  • Solve problems that sit at the cutting edge of AI infrastructure.
  • Build technology with the potential to redefine how humans interact with machines.
  • We're intentionally keeping the team small, talent-dense and highly collaborative.
  • If you want to spend your days solving genuinely difficult problems with exceptional people, we'd love to hear from you.

Robotics is about to have its ChatGPT moment. We're building a new kind of AI system that allows robots to learn complex tasks from a single demonstration. No months of training data collection. No painstaking programming. Show the robot once and it gets to work. Born out of years of research at Imperial College London and backed by one of the largest robotics seed rounds in UK history, we're assembling a small, world-class team to tackle some of the hardest engineering problems in AI. We're looking for an exceptional engineer to own the infrastructure powering our model training. This isn't traditional DevOps. It isn't conventional MLOps. It's a rare opportunity to sit at the intersection of machine learning, distributed systems, cloud infrastructure and robotics.

What you'll do:

Design and own our large-scale AI training infrastructure. Scale distributed model training across hundreds of GPUs and multiple cloud environments. Optimise training throughput, GPU utilisation and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Founding AI Training Infrastructure Architect
Founding AI Training Infrastructure Architect

Imperial College London • Greater London

On-site
GBP 90,000 - 130,000
Equity
Founding-team influence
Small team culture
Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)
Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Good lunch and dinner
Collaborative work culture
No bureaucracy
AI/ ML Infrastructure Engineer
AI/ ML Infrastructure Engineer

OpenSourced - Search & Selection • Bristol

Hybrid
GBP 90,000 - 110,000
Work on real-world AI systems
Direct impact on robotics capability
Fast-moving engineering environment
Principal Machine Learning Infrastructure Engineer London, United Kingdom
Principal Machine Learning Infrastructure Engineer London, United Kingdom

PhysicsX Ltd • Greater London

On-site
GBP 80,000 - 100,000
Equity options
10% employer pension contribution
Free office lunches
+6
Principal Machine Learning Infrastructure Engineer
Principal Machine Learning Infrastructure Engineer

PhysicsX • City Of London

On-site
GBP 80,000 - 120,000
Equity options
10% employer pension contribution
Free office lunches
+2
Software Engineer, Model Deployment- ChatGPT Engineering
Software Engineer, Model Deployment- ChatGPT Engineering

OpenAI • Greater London

On-site
GBP 194,000 - 280,000
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

On-site
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Member of Technical Staff (Post Training)
Member of Technical Staff (Post Training)

Inherentlabs • Greater London

On-site
GBP 80,000 - 100,000
Founding Machine Learning Research Engineer
Founding Machine Learning Research Engineer

Generative • Greater London

On-site
GBP 120,000 - 160,000