AI Infra SRE: Scale Kubernetes, Linux & 24/7 Ops

Radiant

Gloucester

Hybrid

GBP 90,000 - 130,000

Full time

26 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

25 days annual leave
Private medical insurance
Cycle to Work
Gympass access
Stock program

Job summary

Radiant is seeking a senior Site Reliability Engineer in the UK to design, deploy, and operate scalable AI-native infrastructure. You will own Kubernetes clusters, tune Linux and I/O, and drive automation across the platform.

You will champion ITSM practices, maintain Prometheus/Grafana monitoring, and participate in 24x7 on-call support. Mentoring and cross-training with Platform SRE and HPC teams are key parts of the role.

Qualifications

  • 5+ years in globally scaled, 24/7 environments as an SRE or similar role.
  • 3+ years running and optimising orchestration platforms with Kubernetes.
  • Expert Linux administration with Ubuntu distributions.
  • Strong networking (TCP/IP, DNS, DHCP, VLANs).
  • Proficient in Bash and Python automation; ITSM familiarity.

Responsibilities

  • Deploy and manage Kubernetes clusters at scale for AI workloads.
  • Develop manifests and operators for deployment and networking services.
  • Tune Linux kernel, I/O, storage to support orchestration layer workloads.
  • Automate platform lifecycle; build tooling for incident resolution.
  • Apply ITSM processes: Incident, Major Incident, Change Management.
  • Maintain observability stack: Prometheus, Grafana, alerts integrations.
  • Provide 24x7 on-call coverage and incident postmortems.
  • Mentor junior engineers and align with Platform/SRE teams.

Skills

SRE in 24/7 environments
Kubernetes administration
Linux administration (Ubuntu)
Networking fundamentals
Scripting (Bash, Python)
Observability (Prometheus, Grafana)
Incident response & postmortems

Education

Bachelor or Master in Computer Science or related field

Tools

Kubernetes
Prometheus
Grafana
Ansible
Ubuntu Linux

Job description

Radiant is seeking a senior Site Reliability Engineer in the UK to design, deploy, and operate scalable AI-native infrastructure. You will own Kubernetes clusters, tune Linux and I/O, and drive automation across the platform.

You will champion ITSM practices, maintain Prometheus/Grafana monitoring, and participate in 24x7 on-call support. Mentoring and cross-training with Platform SRE and HPC teams are key parts of the role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

24/7 HPC Infra SRE for AI & GPU Compute
24/7 HPC Infra SRE for AI & GPU Compute

Radiant • Gloucester

On-site
GBP 90,000 - 120,000
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability

Radiant • England

On-site
GBP 70,000 - 90,000
Exposure to industry-leading GPU and AI infrastructure
Collaborative, inclusive, and supportive engineering culture
Real ownership and influence over operational excellence
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Radiant • Gloucester

Hybrid
GBP 90,000 - 130,000
25 days annual leave
Private medical insurance
Cycle to Work
+2
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability

Radiant • Greater London

On-site
GBP 90,000 - 140,000
Staff Platform SRE: Scale Infra & Kubernetes Architect
Staff Platform SRE: Scale Infra & Kubernetes Architect

Index Exchange • Greater London

On-site
GBP 120,000 - 160,000
Health benefits
Equity
Parental leave
+3
Senior Platform Engineer (Kubernetes/SRE) – API & Infra Focus
Senior Platform Engineer (Kubernetes/SRE) – API & Infra Focus

FORT • Manchester

On-site
GBP 90,000 - 120,000
Senior DevOps Engineer – Kubernetes & On-Prem AI Infra
Senior DevOps Engineer – Kubernetes & On-Prem AI Infra

Harnham - Data and Analytics Recruitment • Greater London

On-site
GBP 108,000 - 132,000
Salary up to £120,000
Up to 50% bonus
Senior Backend Engineer - Go & Kubernetes for AI Platforms
Senior Backend Engineer - Go & Kubernetes for AI Platforms

Radiant • Greater London

On-site
GBP 50,000 - 95,000
Private medical insurance (Bupa)
Cycle to Work Scheme
Gympass subscription
+3
Senior SRE Architect: Reliability & Self-Healing at Scale
Senior SRE Architect: Reliability & Self-Healing at Scale

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
Scale-Focused AI Infra & MLOps Engineer
Scale-Focused AI Infra & MLOps Engineer

EngineersOfAI • Greater London

On-site
GBP 90,000 - 120,000