Lead HPC Systems Engineer for AI Research Labs

1000scholars

Oxford

On-site

GBP 49,000 - 55,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

University of Oxford seeks a full-time Senior HPC Systems Administrator to lead the development, deployment, and optimization of the lab's HPC infrastructure for AI research in the BOLD Lab.

You will design and maintain CPU/GPU clusters, high-performance storage, and networking, collaborating with researchers, software engineers, industry partners, and IT teams to ensure scalable, secure, and robust systems.

Qualifications

  • Degree in Computer Science, Engineering, or a related technical discipline.
  • Extensive experience managing HPC infrastructure in research or technical environments.
  • Strong Linux systems administration expertise.
  • Experience with GPU servers and high-performance networking technologies.
  • Experience with scripting and automation (Bash, Python).
  • Knowledge of storage systems and backup/archive procedures.
  • Strong troubleshooting and systems integration skills.
  • Excellent written and verbal communication abilities.
  • Ability to work independently and collaboratively in a research environment.

Responsibilities

  • Designing, building, and maintaining HPC clusters and related infrastructure.
  • Managing Linux compute and storage environments.
  • Supporting GPU-enabled research computing systems.
  • Maintaining high-performance storage and backup solutions.
  • Monitoring system performance, security, and availability.
  • Supporting and mentoring researchers, software engineers, and postgraduate students.
  • Developing documentation, training materials, and operational best practices.
  • Collaborating with departmental and university-wide research IT teams.
  • Evaluating and deploying new technologies to support evolving research requirements.

Skills

HPC infrastructure
Linux systems administration
GPU computing
Scripting & automation
High-performance networking
Storage & backup
Documentation & training
Collaboration skills

Education

Bachelor's degree in Computer Science/Engineering

Tools

SLURM or scheduling systems
Containerisation
Cloud platforms
BeeGFS/Lustre

Job description

University of Oxford seeks a full-time Senior HPC Systems Administrator to lead the development, deployment, and optimization of the lab's HPC infrastructure for AI research in the BOLD Lab.

You will design and maintain CPU/GPU clusters, high-performance storage, and networking, collaborating with researchers, software engineers, industry partners, and IT teams to ensure scalable, secure, and robust systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Systems Administrator
Senior HPC Systems Administrator

1000scholars • Oxford

On-site
GBP 49,000 - 55,000
Senior Research Systems Engineer – Linux, HPC & AI
Senior Research Systems Engineer – Linux, HPC & AI

Corehr • Oxford

On-site
GBP 55,000 - 75,000
38 days annual leave
Family leave schemes (up to 26 weeks +
Hybrid/flexible working
+3
Senior Research Systems Administrator — HPC & AI Infra
Senior Research Systems Administrator — HPC & AI Infra

University of Oxford • Oxford

Hybrid
GBP 55,000 - 75,000
Hybrid working
Generous annual leave (38 days)
Pension scheme
+4
Lead Biomedical HPC Engineer – Research Compute & GPUs
Lead Biomedical HPC Engineer – Research Compute & GPUs

MRC Laboratory of Medical Sciences • Greater London

On-site
GBP 58,000 - 68,000
Defined benefit pension
30 days holiday
Cycle to work scheme
+1
HPC-AI Cluster Architect & Automation Lead
HPC-AI Cluster Architect & Automation Lead

NVIDIA • United Kingdom

On-site
GBP 90,000 - 120,000
HPC Research Engineer — Hybrid, Training & Pension
HPC Research Engineer — Hybrid, Training & Pension

The London School of Economics and Political Science (LSE) • Greater London

Hybrid
GBP 43,000 - 52,000
Occupational pension scheme
Generous annual leave
Hybrid working
+1
Senior AI Systems Administrator: GPU HPC Infra Lead
Senior AI Systems Administrator: GPU HPC Infra Lead

CommonAI Holdings Ltd • Cambridge

Remote
GBP 65,000 - 90,000
Pension
Professional development
Networking opportunities
+1
HPC Systems Engineer: Research Compute & AI Workloads
HPC Systems Engineer: Research Compute & AI Workloads

LinuxRecruit • Greater London

On-site
GBP 60,000 - 90,000
Senior HPC Engineer for Biomedical Computation
Senior HPC Engineer for Biomedical Computation

MRC Laboratory of Medical Sciences • City Of London

On-site
GBP 58,000 - 68,000
Pension scheme
Holiday entitlement
Family leave
+4
Senior HPC Research Engineer: Accelerate Science
Senior HPC Research Engineer: Accelerate Science

UKRI International • Greater London

Hybrid
GBP 63,000 - 73,000