Senior HPC Systems Administrator

1000scholars

Oxford

On-site

GBP 49,000 - 55,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

University of Oxford seeks a full-time Senior HPC Systems Administrator to lead the development, deployment, and optimization of the lab's HPC infrastructure for AI research in the BOLD Lab.

You will design and maintain CPU/GPU clusters, high-performance storage, and networking, collaborating with researchers, software engineers, industry partners, and IT teams to ensure scalable, secure, and robust systems.

Qualifications

  • Degree in Computer Science, Engineering, or a related technical discipline.
  • Extensive experience managing HPC infrastructure in research or technical environments.
  • Strong Linux systems administration expertise.
  • Experience with GPU servers and high-performance networking technologies.
  • Experience with scripting and automation (Bash, Python).
  • Knowledge of storage systems and backup/archive procedures.
  • Strong troubleshooting and systems integration skills.
  • Excellent written and verbal communication abilities.
  • Ability to work independently and collaboratively in a research environment.

Responsibilities

  • Designing, building, and maintaining HPC clusters and related infrastructure.
  • Managing Linux compute and storage environments.
  • Supporting GPU-enabled research computing systems.
  • Maintaining high-performance storage and backup solutions.
  • Monitoring system performance, security, and availability.
  • Supporting and mentoring researchers, software engineers, and postgraduate students.
  • Developing documentation, training materials, and operational best practices.
  • Collaborating with departmental and university-wide research IT teams.
  • Evaluating and deploying new technologies to support evolving research requirements.

Skills

HPC infrastructure
Linux systems administration
GPU computing
Scripting & automation
High-performance networking
Storage & backup
Documentation & training
Collaboration skills

Education

Bachelor's degree in Computer Science/Engineering

Tools

SLURM or scheduling systems
Containerisation
Cloud platforms
BeeGFS/Lustre

Job description

Location: Central Oxford

We are seeking a full-time Senior HPC Systems Administrator to join the Department of Engineering Science at the University of Oxford, working within the newly established The British Open-ended Learning and Discovery (BOLD) Lab, a major new UK Research and Innovation (UKRI) funded initiative that will establish a world-leading AI research laboratory at the University of Oxford.

This is an exciting opportunity for an experienced systems professional to lead the development, administration, and strategic evolution of the labs high-performance computing (HPC) infrastructure supporting cutting-edge fundamental AI research, with three initial research pillars: Beyond backpropagation, human centric AI learning and discovery, and embodied learning.

The successful candidate will take a leading role in the design, deployment, maintenance, and optimisation of large-scale HPC systems, including CPU and GPU clusters, high-performance storage, and advanced networking technologies. The postholder will work closely with academic researchers, software engineers, industry partners, and departmental IT teams to ensure robust, scalable, and secure computational infrastructure capable of supporting world-leading research in machine learning and visual computing.

You will possess extensive experience in Linux systems administration and HPC environments, together with strong technical expertise in cluster management, storage systems, networking, scripting, and infrastructure automation. Experience with GPU computing environments, containerisation technologies, cloud platforms, and scientific software environments will be highly desirable. The role also requires excellent communication skills and the ability to collaborate effectively with both technical and non-technical stakeholders.

The role includes responsibility for:

  • Designing, building, and maintaining HPC clusters and associated infrastructure
  • Managing Linux compute and storage environments
  • Supporting GPU-enabled research computing systems
  • Maintaining high-performance storage and backup solutions
  • Monitoring system performance, security, and availability
  • Supporting and mentoring researchers, software engineers, and postgraduate students
  • Developing documentation, training materials, and operational best practices
  • Collaborating with departmental and university-wide research IT teams
  • Evaluating and deploying new technologies to support evolving research requirements

The successful candidate will work under the direction of the departmental HPC Infrastructure Architect, with approximately 20% of their time allocated to broader departmental infrastructure activities.

Essential Skills and Experience
  • Degree in Computer Science, Engineering, or a related technical discipline
  • Extensive experience managing HPC infrastructure in research or technical environments
  • Strong Linux systems administration expertise
  • Experience with GPU servers and high-performance networking technologies
  • Experience with scripting and automation (e.g. Bash, Python)
  • Knowledge of storage systems and backup/archive procedures
  • Strong troubleshooting and systems integration skills
  • Excellent written and verbal communication abilities
  • Ability to work independently and collaboratively in a research environment
Desirable Skills
  • Experience with SLURM or other job scheduling systems
  • Experience with containerisation and virtualisation technologies
  • Knowledge of cloud computing platforms
  • Familiarity with scientific computing tools and AI/ML research workflows
  • Experience supporting research software or academic computing environments
  • Knowledge of high-performance file systems such as BeeGFS or Lustre

The Department holds an Athena Swan Bronze Award, recognising its commitment to advancing gender equality and supporting an inclusive working environment.

  • High Performance Computing (HPC)
  • Linux Systems Administration
  • GPU Computing
  • Cloud & Container Technologies
  • Computer Vision & AI Infrastructure

£49,119 to £55,031 per annum. Grade 8

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Research Systems Administrator
Senior Research Systems Administrator

Corehr • Oxford

On-site
GBP 55,000 - 75,000
38 days annual leave
Family leave schemes (up to 26 weeks +
Hybrid/flexible working
+3
Senior Research Systems Administrator
Senior Research Systems Administrator

University of Oxford • Oxford

On-site
GBP 55,000 - 75,000
Hybrid working
Generous annual leave (38 days)
Pension scheme
+4
Lead HPC Systems Engineer for AI Research Labs
Lead HPC Systems Engineer for AI Research Labs

1000scholars • Oxford

On-site
GBP 49,000 - 55,000
Senior Research HPC Engineer
Senior Research HPC Engineer

UKRI International • Greater London

Hybrid
GBP 63,000 - 73,000
Senior Research HPC Engineer
Senior Research HPC Engineer

UKRI • Greater London

On-site
GBP 58,000 - 68,000
Senior AI Systems Administrator
Senior AI Systems Administrator

CommonAI C.I.C. • Cambridge

On-site
GBP 60,000 - 90,000
Collaborative environment
Growth opportunities
Competitive salary + pension
+3
Senior AI Systems Administrator
Senior AI Systems Administrator

CommonAI Holdings Ltd • Cambridge

Remote
GBP 65,000 - 90,000
Pension
Professional development
Networking opportunities
+1
HPC Support Analyst
HPC Support Analyst

University of Cambridge • Cambridge

On-site
GBP 42,000 - 64,000
36 days holiday per year
Generous pension
Hybrid working
+2
Research Computing Storage Administrator
Research Computing Storage Administrator

Corehr • Oxford

On-site
GBP 52,000 - 76,000
Hybrid and flexible working
Cycle loan scheme
Discounted bus travel
+1
Senior Research Systems Engineer – Linux, HPC & AI
Senior Research Systems Engineer – Linux, HPC & AI

Corehr • Oxford

On-site
GBP 55,000 - 75,000
38 days annual leave
Family leave schemes (up to 26 weeks +
Hybrid/flexible working
+3