Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability

Radiant

England

On-site

GBP 70,000 - 90,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Exposure to industry-leading GPU and AI infrastructure
Collaborative, inclusive, and supportive engineering culture
Real ownership and influence over operational excellence

Job summary

Radiant in the United Kingdom is seeking a Senior Infrastructure Site Reliability Engineer, responsible for ensuring the reliability and performance of high-performance computing infrastructure. This role demands expertise in large-scale distributed systems and operational excellence within a 24/7 support model.

The ideal candidate will have extensive experience with GPU technologies, Linux systems, and performance tuning. Join us to work with advanced technology that influences next-generation compute environments.

Qualifications

  • 8+ years experience in Site Reliability Engineering, Infrastructure Engineering, or similar roles.
  • 2–3+ years recent experience in HPC or AI infrastructure.
  • Strong Linux expertise, especially in production troubleshooting.

Responsibilities

  • Operate and improve high-density AI/HPC infrastructure in a 24/7 production environment.
  • Lead performance evaluation, testing, and operational acceptance of new HPC infrastructure.
  • Drive continuous service improvement through automation and process refinement.

Skills

Site Reliability Engineering
High-performance computing
Linux expertise
Networking fundamentals
Performance tuning
Infrastructure automation
Troubleshooting

Education

Bachelor or Masters Level degree in Computer Science or related field

Tools

NVIDIA GPU ecosystems
Ansible
Prometheus
Grafana
IPMI
iLO

Job description

Radiant in the United Kingdom is seeking a Senior Infrastructure Site Reliability Engineer, responsible for ensuring the reliability and performance of high-performance computing infrastructure. This role demands expertise in large-scale distributed systems and operational excellence within a 24/7 support model.

The ideal candidate will have extensive experience with GPU technologies, Linux systems, and performance tuning. Join us to work with advanced technology that influences next-generation compute environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Infra SRE: GPU Compute, 24/7 Reliability
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability

Radiant • Greater London

On-site
GBP 90,000 - 140,000
24/7 Cloud Infra Support Engineer for AI & HPC
24/7 Cloud Infra Support Engineer for AI & HPC

Radiant • Greater London

On-site
GBP 52,000 - 80,000
25 days annual leave
Private medical insurance (Bupa)
Cycle to Work Scheme
+2
HPC Infrastructure Site Reliability Engineer
HPC Infrastructure Site Reliability Engineer

Radiant • Greater London

On-site
GBP 90,000 - 140,000
Senior Data Center Engineer — GPU & Linux Networking
Senior Data Center Engineer — GPU & Linux Networking

Pursuu • Manchester

Hybrid
GBP 40,000 - 70,000
Company events
Company pension
Free parking
+2
Lead AI GPU Compute Cluster Architect
Lead AI GPU Compute Cluster Architect

Radiant • Greater London

On-site
GBP 120,000 - 190,000
25 days leave
Medical insurance
Cycle to Work
+3
East London HPC Data Centre Operations Engineer
East London HPC Data Centre Operations Engineer

Radiant • Greater London

On-site
GBP 70,000 - 110,000
On-site in East London
Exposure to NVIDIA GPU AI hardware
Global, multi-discipline engineering
Senior Backend Engineer — AI Infra & Kubernetes Leader
Senior Backend Engineer — AI Infra & Kubernetes Leader

Radiant • Greater London

On-site
GBP 110,000 - 150,000
25 days leave
Private medical insurance
Cycle to Work
+3
Senior GPU & AI Infra Architect — Remote, 4-Day Week
Senior GPU & AI Infra Architect — Remote, 4-Day Week

Civo Ltd • United Kingdom

Hybrid
GBP 110,000 - 170,000
4-day week
Uncapped holidays
Remote work environment
Senior HPC Engineer - Hybrid GPU Linux Clusters
Senior HPC Engineer - Hybrid GPU Linux Clusters

Gazelle Global Consulting Limited • Stevenage

Hybrid
GBP 60,000 - 90,000
HPC & AI Platform Engineer – GPU Clusters
HPC & AI Platform Engineer – GPU Clusters

Era4 • United Kingdom

Hybrid
GBP 95,000 - 130,000