AI Supercomputing Infrastructure Associate Engineer

University of Bristol

West of England

On-site

GBP 65,000 - 90,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

The University of Bristol's AI Supercomputing Centre seeks a capable engineer to build and operate software-defined supercomputing platforms for researchers. You will design and maintain large-scale GPU/CPU/GPU workloads, collaborate with researchers, and help scale services using Kubernetes, Terraform, and Python/Rust.

The role focuses on robust service delivery, open access infrastructure, and distributed systems at national-scale research within cybersecurity-compliant practices.

Qualifications

  • Experience with large distributed systems and infrastructure as code.
  • Strong collaboration with researchers and operations teams.

Responsibilities

  • Use tools such as Python, Rust, Terraform/OpenTofu, Kubernetes, Git and Bash.
  • Design and operate large, highly available supercomputing services managed as software-defined infrastructures.
  • Experience designing and operating GPU and CPU/GPU workloads at scale.
  • Co-design solutions with researchers to enable new algorithms and software.

Skills

SysOps
NetOps
DevOps
SecOps
MLOps
Research Software Engineering

Education

Degree in Computer Science or related field
Equivalent practical experience in ML/AI research

Tools

Python
Rust
Terraform/OpenTofu
Kubernetes
Git
Bash

Job description

The Bristol AI Supercomputing Centre runs the Isambard-AI National Artificial Intelligence Research Resource, recently announced AI Data facility and the Isambard3 Tier-2 Supercomputer. Isambard-AI is the most powerful supercomputer in the UK and amongst the most powerful in Europe.

  • The AI Supercomputing team owns the entire process of developing and operating the centre's compute and software infrastructure, which includes:
  • The sourcing of hardware and system design.
  • The deployment of huge software-defined infrastructure using tools such as Kubernetes and Terraform / OpenTofu.
  • Building and operating platforms to enable researchers to conduct leading-edge research using the systems.
  • Optimising and refining software to ensure environmental and economic efficient use.

As one of the largest Open AI Research Resources internationally, we are committed to catalysing an AI transformation in the research and development community. In this role, you will work as part of the AI Supercomputing Team to build and operate primarily the infrastructure and compute platforms that researchers use for their work. You do not need to be an AI or computational research domain expert to deliver world-class infrastructure, but you do need to quickly obtain a deep technical understanding of new domains. You should enjoy being self-directed and identifying the most important problems to solve as the team matures with standardized tools and processes around stability, robust service delivery and scaling.

What will you be doing?
  • Use tools such as Python, Rust, Terraform / OpenTofu, Kubernetes, Git and Bash.
  • Design and operate large, highly available supercomputing services managed as software-defined infrastructures, and integrated as complete computational experiments.
  • You will experience designing and operating massive-scale GPU and combined CPU/GPU workloads across these services.
  • You will design and debug platforms, and will work closely with researchers as you co-design solutions that will enable the development and operation of new algorithms and software to solve leading-edge research problems.
You should apply if
  • Want to help build, maintain and secure some of the largest, modern software-defined supercomputing systems.
  • Would enjoy working with world class domain and AI researchers as your primary workload .
  • Have supported development or operation of a small to large clusters or dabbled in building your own physical or software-defined systems.
  • Love operating large distributed, highly available systems, and want to see them used for truly open national-scale research in a cybersecurity compliant manner.
  • Domain knowledge in 1 or more areas from SysOps, NetOps, DevOps, SecOps, MLOps or Research Software Engineering.
  • Degree (or equivalent practical experience) in computer science, computational or ML/AI research or in a natural science with a high degree of competence in computer science or computational research.
  • Ability to contribute as a member of diverse, technical teams and to follow operational procedures.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Associate Engineer (3 x FTE)
AI Infrastructure Associate Engineer (3 x FTE)

University of Bristol • West of England

Hybrid
GBP 45,000 - 62,000
Associate AI Supercomputing Infrastructure Engineer
Associate AI Supercomputing Infrastructure Engineer

University of Bristol • West of England

On-site
GBP 65,000 - 90,000
AI Infrastructure Engineer – Storage & Networking
AI Infrastructure Engineer – Storage & Networking

University of Bristol • West of England

Hybrid
GBP 45,000 - 62,000
Senior Systems Engineer - Commissioning and Characterisation
Senior Systems Engineer - Commissioning and Characterisation

United States Digital Space LLC • West of England, Greater London

On-site
GBP 70,000 - 110,000
Flexible working
Generous leave
Pension matching
+2
Systems Research Engineer (AI Infrastructure & Distributed Systems
Systems Research Engineer (AI Infrastructure & Distributed Systems

European Tech Recruit • City of Edinburgh

On-site
GBP 60,000 - 80,000
HPC Support Analyst
HPC Support Analyst

University of Cambridge • Cambridge

Hybrid
GBP 42,000 - 64,000
36 days holiday per year
Generous pension
Hybrid working
+2
Cyber Security Engineer - Cyber & Autonomous Systems Team
Cyber Security Engineer - Cyber & Autonomous Systems Team

AI Security Institute • City Of London

On-site
GBP 70,000 - 110,000
Staff Engineer, Digital Infrastructure (R5602)
Staff Engineer, Digital Infrastructure (R5602)

Shield AI • Greater London

On-site
GBP 95,000 - 135,000
Senior AI Software Engineer
Senior AI Software Engineer

Ocho People • Belfast City District

On-site
GBP 90,000 - 120,000
Share options
AI tooling culture
Misuse Red Team - Research Engineer/Research Scientist
Misuse Red Team - Research Engineer/Research Scientist

AI Security Institute • Greater London

On-site
GBP 65,000 - 145,000
Hybrid working
Pension contribution
Conference funding
+1