Senior HPC Infra SRE: GPU Compute, 24/7 Reliability

Radiant

Greater London

On-site

GBP 90,000 - 140,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Radiant is seeking a senior Infrastructure Site Reliability Engineer for its GPU-accelerated HPC platforms. You will own reliability and performance across large-scale distributed systems, operating in a 24/7 on-call environment.

The role demands deep Linux administration, extensive HPC experience, and strong collaboration with Platform, Network, and Data Centre teams to improve observability and drive automation at scale.

Qualifications

  • 8+ years in Site Reliability or similar roles.
  • 2–3+ years in HPC/AI infra.
  • Strong Linux administration and troubleshooting.
  • Experience with ITIL-aligned processes.
  • Experience with GPU/InfiniBand/NVIDIA ecosystems.

Responsibilities

  • Operate and improve 24/7 HPC infrastructure.
  • Lead on-call rotation and incident response.
  • Drive CSI and automation for reliability.
  • Collaborate across Platform, Network, and Data Centre teams.
  • Contribute to future HPC design.

Skills

SRE experience
HPC/AI infra
Linux systems
Performance tuning
Bare-metal infra
Networking fundamentals
NVIDIA GPUs
Ansible/Python
Observability
ITIL processes

Education

Bachelor or Masters in CS/Engineering

Tools

IPMI/iLO/iDRAC
Redfish
Prometheus/Grafana
Kubernetes exposure

Job description

Radiant is seeking a senior Infrastructure Site Reliability Engineer for its GPU-accelerated HPC platforms. You will own reliability and performance across large-scale distributed systems, operating in a 24/7 on-call environment.

The role demands deep Linux administration, extensive HPC experience, and strong collaboration with Platform, Network, and Data Centre teams to improve observability and drive automation at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability

Radiant • England

On-site
GBP 70,000 - 90,000
Exposure to industry-leading GPU and AI infrastructure
Collaborative, inclusive, and supportive engineering culture
Real ownership and influence over operational excellence
HPC Infrastructure Site Reliability Engineer
HPC Infrastructure Site Reliability Engineer

Radiant • Greater London

On-site
GBP 90,000 - 140,000
24/7 Cloud Infra Support Engineer for AI & HPC
24/7 Cloud Infra Support Engineer for AI & HPC

Radiant • Greater London

On-site
GBP 52,000 - 80,000
25 days annual leave
Private medical insurance (Bupa)
Cycle to Work Scheme
+2
GPU HPC Operations Manager - Reliability & Agile Leader
GPU HPC Operations Manager - Reliability & Agile Leader

Northern Data Group • Greater London

Hybrid
GBP 90,000 - 130,000
Platform Observability & Automation Engineer
Platform Observability & Automation Engineer

Radiant • Greater London

On-site
GBP 90,000 - 120,000
Senior Data Center Engineer — GPU & Linux Networking
Senior Data Center Engineer — GPU & Linux Networking

Pursuu • Manchester

Hybrid
GBP 40,000 - 70,000
Company events
Company pension
Free parking
+2
Datacentre Operations Engineer
Datacentre Operations Engineer

Radiant • Greater London

On-site
GBP 70,000 - 110,000
On-site in East London
Exposure to NVIDIA GPU AI hardware
Global, multi-discipline engineering
Senior HPC Engineer
Senior HPC Engineer

Gazelle Global Consulting Limited • Stevenage

Hybrid
GBP 60,000 - 90,000
Senior HPC Engineer - Hybrid GPU Linux Clusters
Senior HPC Engineer - Hybrid GPU Linux Clusters

Gazelle Global Consulting Limited • Stevenage

Hybrid
GBP 60,000 - 90,000
Senior HPC Platform Engineer — Reliability & Automation
Senior HPC Platform Engineer — Reliability & Automation

Mercedes AMG High Performance Powertrains • Brixworth

On-site
GBP 65,000 - 90,000