HPC Infrastructure Site Reliability Engineer

Radiant

Gloucester

On-site

GBP 90,000 - 120,000

Full time

25 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Radiant is seeking a senior Infrastructure Site Reliability Engineer to own and improve large‑scale GPU‑accelerated HPC infrastructure in a 24/7 production environment.

You will work across network, storage, virtualization and orchestration with hands‑on Linux expertise, NVIDIA GPU ecosystems, RoCE/InfiniBand, and performance benchmarking.

This role champions observability, automation and on‑call reliability, shaping next‑gen HPC platforms within a globally distributed team.

Qualifications

  • Experience operating large-scale distributed HPC environments.
  • Strong Linux and system troubleshooting skills.
  • Hands-on experience with GPU-accelerated HPC and AI infrastructure.

Responsibilities

  • Operate and improve high-density AI/HPC infrastructure in a 24/7 production environment.
  • Participate in a 24x7x365 on-call rotation.
  • Drive CSI initiatives to reduce toil and improve reliability and observability.

Skills

GPU HPC
Linux
On‑call experience
Distributed systems
Networking
Performance benchmarking

Tools

NVIDIA GPU ecosystems
RoCE/InfiniBand
Virtualization
Orchestration

Job description

About Us

We’re a fast-growing GPU-as-a-Service provider, delivering scalable, high-performance compute infrastructure purpose‑built for AI and HPC workloads. Operating across global data centres, we run mission‑critical environments where uptime, throughput, and ultra‑low latency are non‑negotiable.


About Us

We’re a fast-growing GPU-as-a-Service provider, delivering scalable, high-performance compute infrastructure purpose‑built for AI and HPC workloads. Operating across global data centres, we run mission‑critical environments where uptime, throughput, and ultra‑low latency are non‑negotiable.


Role Overview

We are looking for a senior Infrastructure Site Reliability Engineer with deep experience operating large‑scale distributed systems and recent hands‑on expertise in high‑performance computing (HPC) and AI infrastructure. This is an operations‑first SRE role, working in a 24/7/365 on‑call environment, responsible for ensuring reliability, performance, and continuous improvement of mission‑critical infrastructure. This role sits within a cross‑functional organisation spanning network engineering, infrastructure SRE, Platform SRE, infrastructure tooling engineers (software) and data centre operations.


The ideal candidate has progressed through large‑scale, globally distributed or multi‑site infrastructure environments and has more recently specialised in GPU‑accelerated HPC systems. This role provides exposure to the latest high‑density AI compute platforms, including next‑generation GPU infrastructure at significant scale. You will bring strong breadth across bare metal, networking, storage, virtualisation, and orchestration, alongside deep HPC experience including NVIDIA GPU ecosystems, RDMA networking (RoCE and InfiniBand), and performance validation and benchmarking. Strong Linux and distributed systems expertise is essential.


Alongside operational ownership, this is a deeply technical Infrastructure SRE role centred on advanced operational troubleshooting and performance evaluation across large‑scale HPC systems. You will investigate complex, cross‑layer issues spanning GPU compute, networking, storage, and orchestration, building a clear understanding of system behaviour under real production AI and HPC workloads.


A key responsibility is performance evaluation, testing, and operational acceptance of new HPC environments, ensuring platforms meet defined reliability, scalability, and performance expectations before entering production. You will work across hardware, network, and software layers to validate readiness of high‑density GPU infrastructure and support safe, predictable deployment at scale.


You will also play a central role in continuous service improvement (CSI)—reducing operational toil, increasing automation, and improving reliability, consistency, and operational efficiency across the platform. This includes strengthening observability, refining operational workflows, and eliminating repetitive or failure‑prone processes.


Over time, you will help shape future infrastructure design and deployment approaches, feeding operational insight back into infrastructure engineering decisions and ensuring production learnings directly influence next‑generation HPC platform evolution.


What’s In It For You

Join a team operating some of the world’s most advanced high‑performance computing infrastructure. As a HPC Infrastructure SRE, you’ll work hands‑on with cutting‑edge GPU and CPU platforms ‑ including the latest NVIDIA architectures ‑ powering dense, large‑scale compute environments used for AI, machine learning, and next‑generation workloads.


This is an opportunity to build expertise at the forefront of modern infrastructure, where reliability, scale, and performance matter every day. You’ll collaborate with experienced engineers across a globally distributed organisation that values openness, inclusion, technical excellence, and continuous learning.


We move quickly, solve meaningful challenges, and give people the space to make an impact. If you thrive in fast‑paced environments, enjoy working with advanced technology, and want to help shape the future of high‑performance compute, you’ll find both challenge and opportunity here.


You can also expect:


  • Exposure to industry‑leading GPU and AI infrastructure

  • Opportunities to grow alongside a rapidly scaling global business

  • A collaborative, inclusive, and supportive engineering culture

  • Real ownership and the ability to influence operational excellence

  • Work that sits at the intersection of people, performance, and technology

  • A modern, flexible, globally connected workplace with ambitious goals

  • Key Responsibilities

  • Operate and improve high‑density AI/HPC infrastructure in a 24/7 production environment

  • Participate in a 24x7x365 on‑call rotation, supporting mission
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Infrastructure Site Reliability Engineer
HPC Infrastructure Site Reliability Engineer

Radiant • Greater London

On-site
GBP 90,000 - 140,000
Datacentre Operations Engineer
Datacentre Operations Engineer

Radiant • Greater London

On-site
GBP 70,000 - 110,000
On-site in East London
Exposure to NVIDIA GPU AI hardware
Global, multi-discipline engineering
24/7 HPC Infra SRE for AI & GPU Compute
24/7 HPC Infra SRE for AI & GPU Compute

Radiant • Gloucester

On-site
GBP 90,000 - 120,000
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability

Radiant • Greater London

On-site
GBP 90,000 - 140,000
Operations Engineering Manager (m/f/d)
Operations Engineering Manager (m/f/d)

Northern Data Group • Greater London

Hybrid
GBP 90,000 - 130,000
Infrastructure Tooling & Observability Engineer( UK)
Infrastructure Tooling & Observability Engineer( UK)

Radiant • Greater London

On-site
GBP 90,000 - 120,000
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability

Radiant • England

On-site
GBP 70,000 - 90,000
Exposure to industry-leading GPU and AI infrastructure
Collaborative, inclusive, and supportive engineering culture
Real ownership and influence over operational excellence
VP Data Centre Operations
VP Data Centre Operations

WNTD • Greater London

On-site
GBP 100,000 - 150,000
Platform Engineer
Platform Engineer

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Incite-Insight.co.uk • West of England

On-site
GBP 70,000 - 95,000