Senior ML Infra Engineer: Scale GPU Clusters & Pipelines

Ellison Institute, LLC

Oxford

Hybrid

GBP 90,000 - 140,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive salary
25 days annual leave + 8 bank holidays
3 additional days between Christmas &.
Pension - Employer contribution 7.5%
Life Assurance
Income Protection
Private Medical Insurance for you and/
Employee discounts
Electric car scheme
Nursery Salary Sacrifice scheme
Cycle to Work Scheme
Family Planning
Neurodiversity support
Coaching & Therapy services

Job summary

Ellison Institute, LLC is seeking an experienced Senior ML infrastructure engineer to build and operate scalable GPU training and inference clusters in a secure, observable environment. You will design high-throughput data paths and manage lifecycle of compute and storage resources to accelerate research outcomes.

Ideal candidates bring hands-on experience with containerised ML pipelines, GPUs, and modern CI/CD practices.

Qualifications

  • Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale.
  • Autonomous approach to systems design with ability to ideate, co-create and implement optimal solutions.
  • Experience migrating ML infrastructure from traditional schedulers to modern containerised systems.
  • Expertise with high-throughput storage systems for ML/HPC workloads.
  • Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling.

Responsibilities

  • Build, operate, and continuously optimise our high-performance GPU training and inference clusters, focusing on robust, high-availability scheduling, isolation, and automated lifecycle management.
  • Drive systems design and implementation for high-throughput data paths, optimising I/O, caching, and data locality across compute and storage (including Lustre).
  • Proactively benchmark, profile, and resolve performance bottlenecks across the compute, network, and orchestration layers to maximise efficiency for distributed training and inference.
  • Establish comprehensive observability, resilience, and automated security controls to ensure compliance and robust operation of sensitive research environments.
  • Partner with Research, Data, and Applied teams to forecast capacity and cost for GPU and storage needs, setting quotas and streamlining ML experimentation pipelines.

Skills

ML compute clusters
Systems design
Containerised infra
GPU networking
CI/CD (Terraform, Argo CD)

Tools

Kubernetes
Terraform
Argo CD

Job description

Ellison Institute, LLC is seeking an experienced Senior ML infrastructure engineer to build and operate scalable GPU training and inference clusters in a secure, observable environment. You will design high-throughput data paths and manage lifecycle of compute and storage resources to accelerate research outcomes.

Ideal candidates bring hands-on experience with containerised ML pipelines, GPUs, and modern CI/CD practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Compute Architect: Scalable GPU Pipelines
Senior ML Compute Architect: Scalable GPU Pipelines

Ellison Institute of Technology • Oxford

On-site
GBP 90,000 - 130,000
Travel allowance
Pension 7.5%
Private Medical Insurance
+2
ML Performance Engineer – Scale GPU/CPU Workloads
ML Performance Engineer – Scale GPU/CPU Workloads

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
Senior Cloud Platform Engineer — AI, GPU Compute & Secure CI/CD
Senior Cloud Platform Engineer — AI, GPU Compute & Secure CI/CD

Ellison Institute, LLC • Oxford

Hybrid
GBP 90,000 - 120,000
Senior Compute Infrastructure Engineer – AI Training & LLMs
Senior Compute Infrastructure Engineer – AI Training & LLMs

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Good lunch and dinner
Collaborative work culture
No bureaucracy
ML Performance Engineer: Scale & Optimize GPU/CPU Workloads
ML Performance Engineer: Scale & Optimize GPU/CPU Workloads

G-Research • Greater London

On-site
GBP 90,000 - 135,000
Lunch provided (Just Eat for Business)
Barista bar
35 days annual leave
+5
Lead GPU Infrastructure Architect for Scalable AI Clusters
Lead GPU Infrastructure Architect for Scalable AI Clusters

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
Senior ML Platform Engineer - Scalable AI Deployment
Senior ML Platform Engineer - Scalable AI Deployment

Scale AI, Inc. • Greater London

On-site
GBP 100,000 - 150,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • Greater London

On-site
GBP 65,000 - 90,000
Machine Learning & Cloud Infra Engineer
Machine Learning & Cloud Infra Engineer

SpAItial AI • Greater London

On-site
GBP 60,000 - 85,000
Senior ML Engineer: GPU Inference & Low-Precision Training
Senior ML Engineer: GPU Inference & Low-Precision Training

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive pay
Career growth
Flexibility and ownership
+3