Senior HPC/GPU Infra SRE - 24x7 On-Call

Radiant

United Kingdom

On-site

GBP 90,000 - 150,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Radiant is building out a senior Infrastructure Site Reliability Engineer role focused on large-scale GPU‑accelerated HPC infrastructure. The role is operations‑first, in a 24/7 on‑call environment, spanning bare metal, networking, storage, virtualization, and orchestration with deep HPC expertise.

You will improve reliability, performance, and observability, shaping future HPC platform deployments and driving CSI across the stack.

Qualifications

  • 8+ years in SRE or infrastructure in large-scale, 24/7 environments.
  • 2-3+ years HPC/AI infrastructure experience with GPU compute.
  • Strong Linux production troubleshooting and tuning skills.
  • Experience with InfiniBand/RoCE networks and HPC fabrics.
  • IaC and scripting experience (Bash, Python, Ansible, etc.).

Responsibilities

  • Operate and improve high-density AI/HPC infrastructure in a 24/7 production environment.
  • Lead 24x7x365 on-call rotation, incident response, and RCA follow-up.
  • Troubleshoot cross-layer issues across compute, networking, storage, and orchestration.
  • Lead performance evaluation and acceptance of new HPC infrastructure before production.
  • Drive continuous service improvement by automation, tooling, and process refinement.
  • Build and maintain infrastructure automation and observability tooling at scale.
  • Collaborate with Platform SRE, Network, and Data Centre Ops to enhance reliability.

Skills

Site Reliability Engineering
Infrastructure Engineering
Linux (Ubuntu)
Networking fundamentals
Bare-metal hardware
IPMI / iLO / iDRAC
Prometheus / Grafana
Python / Bash / Ansible

Education

Bachelor's degree in Computer Science, Engineering or related field
LPIC Certifications

Tools

IPMI
iLO
iDRAC
Redfish
Kubernetes exposure

Job description

Radiant is building out a senior Infrastructure Site Reliability Engineer role focused on large-scale GPU‑accelerated HPC infrastructure. The role is operations‑first, in a 24/7 on‑call environment, spanning bare metal, networking, storage, virtualization, and orchestration with deep HPC expertise.

You will improve reliability, performance, and observability, shaping future HPC platform deployments and driving CSI across the stack.

Get your free, confidential resume review.
or drag and drop your file here.