Infrastructure Site Reliability Engineer

Radiant

Gloucester

On-site

GBP 70,000 - 110,000

Full time

38 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

25 days annual leave
Cycle to Work Scheme
Gympass subscription

Job summary

Radiant is pursuing an experienced Infrastructure Site Reliability Engineer to run and evolve our AI-native infrastructure stack in the UK. You’ll cover bare-metal, virtualization, and orchestration layers while mentoring teammates and improving automation to support AI/HPC workloads.

You will configure and operate resilient Linux systems (Ubuntu), refine performance, and contribute to the observability stack with Prometheus and Grafana.

Qualifications

  • Experience operating in 24/7 performance‑critical environments.
  • Expert Linux administration, Ubuntu preferred.
  • Strong networking and automation scripting skills.
  • Familiarity with ITSM and incident response is a bonus.

Responsibilities

  • Deploy and operate scalable infrastructure for AI/HPC workloads.
  • Configure and maintain bare-metal and virtualization layers.
  • Develop automation scripts and IaC to support platform lifecycle.
  • Lead on-call rotation and incident response processes.

Skills

Linux administration
Networking fundamentals
Automation scripting
Observability tooling
Kubernetes
Scripting: Bash
Python

Education

Bachelor/Master in CS/Engineering or related

Tools

IPMI
Redfish
PXE
Prometheus
Grafana
MAAS
Tinkerbell

Job description

About Radiant

Radiant is redefining how AI infrastructure is built. We design and operate AI-native cloud platforms engineered for sovereignty, performance, and scale. Our infrastructure powers GPU-native workloads, multi-tenant control planes, and high-performance AI systems designed for the most demanding environments. We are not building a generic cloud. We are building purpose-built AI infrastructure - from powered land, to compute, to software . As we scale our platform and expand our engineering organisation, we are looking for leaders who can build strong teams, uphold high standards, and deliver reliably at pace.

Job Summary

We’re looking for an experienced Infrastructure Site Reliability Engineer to run and evolve our infrastructure stack. You’ll contribute across bare-metal, virtualization, and orchestration layers, keeping things stable and secure 24/7 x 365 — all while mentoring teammates, improving process and automation as well as helping translate deep technical concepts for a wide range of collaborators and customers.

What You’ll Do
  • Deploy and operate resilient, scalable infrastructure supporting AI/HPC workloads
  • Optimize Linux system configuration, BIOS/firmware, kernel, and disk subsystem for performance
  • Configure, monitor and manage bare-metal infrastructure using IPMI, Redfish, etc
  • Build and maintain automation scripts and infrastructure as code to support platform lifecycle, as well as simplifying troubleshooting for Incident resolution and provision of tooling for our support organisation
  • Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement.
  • Maintain and enhance ’s observability stack: Prometheus, Grafana, and custom monitoring integrations
  • Operate and support services in 24x7 production environments, including on-call rotation
  • Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations
  • Mentor junior engineers and act as an Operational requirements consultant to other departments
  • Communicate technical decisions clearly to non-technical stakeholders and customers
  • Uphold a culture of: do, document, automate
  • Willingness to cross train with Platform Engineering/Platform SRE to fully support both our infrastructure and platform stacks.
  • Willingness to cross train with HPC Engineering, supported by NVIDIA to enhance our
  • HPC supportability offering
What you bring
  • 5+ Years Proven experience in globally scaled, performance-intensive environments operating to a 24/7 support model
  • Expert-level Linux administration, especially Ubuntu distributions
  • Proficiency in system tuning, disk I/O optimization, and hardware-level performance tweaks
  • Familiarity with Out of Band management tools (IPMI, Redfish, PXE, etc.)
  • Strong networking fundamentals: TCP/IP, DNS, DHCP, VLANs, routing, switching
  • Strong experience with infrastructure scripting and automation (Bash, Python, Ansible)
  • Deep understanding of observability principles and tools (Prometheus, Grafana)
  • Hands-on experience operating orchestration platforms (Kubernetes, MAAS, Tinkerbell)
  • Strong grasp of ITSM and service operation best practices
  • Excellent communication and mentorship skills
  • Comfortable interfacing with internal stakeholders and external customers
  • Bonus: Knowledge of HPC workloads and GPU-based infrastructure
  • Bonus: Experience with InfiniBand networks and HPC performance tuning
Nice to have
  • Bachelor or Masters Level degree in Computer Science, Engineering or related field, or equivalent experience.
  • LPIC Certifications
  • ITIL Foundation level qualification or equivalent experience
How you work
  • You approach problems with a systems mindset - balancing practical execution with long-term scalability
  • You elevate the team, setting high standards for technical quality and engineering excellence.
  • You hold yourself and others accountable - giving direct feedback and expecting the same
  • You take initiative, owning challenges end-to-end and proactively driving solutions.
  • You invest in others, mentoring to build both capability and confidence.
  • You communicate clearly - translating complexity into clarity across engineering and business audiences
Why should you join us?

What sets us apart is our blend of modern technology, competitive benefits, and an open, welcoming work culture that enables our people to thrive.

  • 25 days of annual leave
  • A culture that emphasises results over hierarchy, process & ego: we place great emphasis on the quality, ingenuity and creativity of work.
  • Open communication, regular feedback: we value smooth collaboration, direct and actionable feedback, and believe that leading with empathy and a growth mindset makes us better together.
  • Learning Time: we all have dedicated learning time to focus on new skills, projects or interests that lay outside of your day-to-day job.
  • Health & Wellbeing: we want everyone to feel healthy and happy, so we offer private medical insurance via Bupa.
  • Cycle to Work Scheme: we're committed to building a sustainable business, so we encourage cycling to work.
  • Gympass subscription to a variety of gyms and wellbeing apps
  • Participation in the company shares program
  • Enhanced parental pay & leave
Diversity, Equality, Inclusion and Belonging

We are an equal opportunity employer and we strive to reduce unconscious bias throughout our hiring process. All applicants will be considered for employment without attention to ethnicity, religion, sexual orientation, gender identity, family or parental status, national origin, veteran, neurodiversity status or disability status. To ensure our recruitment processes provide an equal opportunity for all applicants to succeed, we encourage you to let us know if there are any adjustments that we can make.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Site Reliability Engineer
Platform Site Reliability Engineer

Radiant • Gloucester

Hybrid
GBP 90,000 - 130,000
25 days annual leave
Private medical insurance
Cycle to Work
+2
Cluster Architect
Cluster Architect

Radiant • Greater London

On-site
GBP 120,000 - 190,000
25 days leave
Medical insurance
Cycle to Work
+3
HPC Infrastructure Site Reliability Engineer
HPC Infrastructure Site Reliability Engineer

Radiant • Greater London

On-site
GBP 90,000 - 140,000
Senior Product Manager - Storage & Networking
Senior Product Manager - Storage & Networking

Jackalope Digital LLC • Greater London

Hybrid
GBP 110,000 - 140,000
25 days annual leave
Health & wellbeing: private medical
Cycle to Work Scheme
+4
Senior Product Manager - Storage & Networking
Senior Product Manager - Storage & Networking

Radiant • Greater London

On-site
GBP 90,000 - 130,000
25 days annual leave
Health and wellbeing benefits
Gympass access
+3
IT Lead Engineer
IT Lead Engineer

PhysicsX • Greater London

On-site
GBP 90,000 - 140,000
Equity options
10% employer pension contribution
Free office lunches
+10
HPC Infrastructure Site Reliability Engineer
HPC Infrastructure Site Reliability Engineer

Radiant • Gloucester

On-site
GBP 90,000 - 120,000
Network Engineer
Network Engineer

asobbi • United Kingdom

Remote
GBP 53,000 - 69,000
Highly competitive package with equity
Dynamic progression plan
Human-first flexibility
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Armstrong Talent Partners • Blantyre

On-site
GBP 60,000 - 90,000
33 days annual leave
Life assurance
Group income protection
+4
Senior Backend Software Engineer - Core Services
Senior Backend Software Engineer - Core Services

Radiant • Greater London

On-site
GBP 50,000 - 95,000
Private medical insurance (Bupa)
Cycle to Work Scheme
Gympass subscription
+3