Lead AI Infra SRE: Scale, Reliability & Mentorship

Radiant

Gloucester

On-site

GBP 70,000 - 110,000

Full time

25 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

25 days annual leave
Cycle to Work Scheme
Gympass subscription

Job summary

Radiant is pursuing an experienced Infrastructure Site Reliability Engineer to run and evolve our AI-native infrastructure stack in the UK. You’ll cover bare-metal, virtualization, and orchestration layers while mentoring teammates and improving automation to support AI/HPC workloads.

You will configure and operate resilient Linux systems (Ubuntu), refine performance, and contribute to the observability stack with Prometheus and Grafana.

Qualifications

  • Experience operating in 24/7 performance‑critical environments.
  • Expert Linux administration, Ubuntu preferred.
  • Strong networking and automation scripting skills.
  • Familiarity with ITSM and incident response is a bonus.

Responsibilities

  • Deploy and operate scalable infrastructure for AI/HPC workloads.
  • Configure and maintain bare-metal and virtualization layers.
  • Develop automation scripts and IaC to support platform lifecycle.
  • Lead on-call rotation and incident response processes.

Skills

Linux administration
Networking fundamentals
Automation scripting
Observability tooling
Kubernetes
Scripting: Bash
Python

Education

Bachelor/Master in CS/Engineering or related

Tools

IPMI
Redfish
PXE
Prometheus
Grafana
MAAS
Tinkerbell

Job description

Radiant is pursuing an experienced Infrastructure Site Reliability Engineer to run and evolve our AI-native infrastructure stack in the UK. You’ll cover bare-metal, virtualization, and orchestration layers while mentoring teammates and improving automation to support AI/HPC workloads.

You will configure and operate resilient Linux systems (Ubuntu), refine performance, and contribute to the observability stack with Prometheus and Grafana.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra SRE: Scale Kubernetes, Linux & 24/7 Ops
AI Infra SRE: Scale Kubernetes, Linux & 24/7 Ops

Radiant • Gloucester

Hybrid
GBP 90,000 - 130,000
25 days annual leave
Private medical insurance
Cycle to Work
+2
Infrastructure Site Reliability Engineer
Infrastructure Site Reliability Engineer

Radiant • Gloucester

On-site
GBP 70,000 - 110,000
25 days annual leave
Cycle to Work Scheme
Gympass subscription
24/7 HPC Infra SRE for AI & GPU Compute
24/7 HPC Infra SRE for AI & GPU Compute

Radiant • Gloucester

On-site
GBP 90,000 - 120,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Radiant • Gloucester

Hybrid
GBP 90,000 - 130,000
25 days annual leave
Private medical insurance
Cycle to Work
+2
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability

Radiant • England

On-site
GBP 70,000 - 90,000
Exposure to industry-leading GPU and AI infrastructure
Collaborative, inclusive, and supportive engineering culture
Real ownership and influence over operational excellence
Senior Cloud SRE: Scale AI Platform & Reliability
Senior Cloud SRE: Scale AI Platform & Reliability

Mistral AI • Greater London

On-site
GBP 75,000 - 110,000
Healthcare coverage
Relocation support
Retirement plans
+3
Scale-Focused AI Infra & MLOps Engineer
Scale-Focused AI Infra & MLOps Engineer

EngineersOfAI • Greater London

On-site
GBP 90,000 - 120,000
Senior Cloud SRE for AI Platform — Reliability & Scale
Senior Cloud SRE for AI Platform — Reliability & Scale

Mistral • Greater London

On-site
GBP 90,000 - 140,000
Healthcare coverage
Parental leave
Retirement plans
+3
Senior SRE - AI Inference Platform, Scale & Reliability
Senior SRE - AI Inference Platform, Scale & Reliability

Nebius • Greater London

On-site
GBP 100,000 - 140,000
Competitive compensation
Career growth and learning opportunity
Flexibility and ownership
+3
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability

Radiant • Greater London

On-site
GBP 90,000 - 140,000