Platform Site Reliability Engineer

Radiant

Gloucester

Hybrid

GBP 90,000 - 130,000

Full time

3 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

25 days annual leave
Private medical insurance
Cycle to Work
Gympass access
Stock program

Job summary

Radiant is seeking a senior Site Reliability Engineer in the UK to design, deploy, and operate scalable AI-native infrastructure. You will own Kubernetes clusters, tune Linux and I/O, and drive automation across the platform.

You will champion ITSM practices, maintain Prometheus/Grafana monitoring, and participate in 24x7 on-call support. Mentoring and cross-training with Platform SRE and HPC teams are key parts of the role.

Qualifications

  • 5+ years in globally scaled, 24/7 environments as an SRE or similar role.
  • 3+ years running and optimising orchestration platforms with Kubernetes.
  • Expert Linux administration with Ubuntu distributions.
  • Strong networking (TCP/IP, DNS, DHCP, VLANs).
  • Proficient in Bash and Python automation; ITSM familiarity.

Responsibilities

  • Deploy and manage Kubernetes clusters at scale for AI workloads.
  • Develop manifests and operators for deployment and networking services.
  • Tune Linux kernel, I/O, storage to support orchestration layer workloads.
  • Automate platform lifecycle; build tooling for incident resolution.
  • Apply ITSM processes: Incident, Major Incident, Change Management.
  • Maintain observability stack: Prometheus, Grafana, alerts integrations.
  • Provide 24x7 on-call coverage and incident postmortems.
  • Mentor junior engineers and align with Platform/SRE teams.

Skills

SRE in 24/7 environments
Kubernetes administration
Linux administration (Ubuntu)
Networking fundamentals
Scripting (Bash, Python)
Observability (Prometheus, Grafana)
Incident response & postmortems

Education

Bachelor or Master in Computer Science or related field

Tools

Kubernetes
Prometheus
Grafana
Ansible
Ubuntu Linux

Job description

About Radiant

Radiant is redefining how AI infrastructure is built.

About Radiant

Radiant is redefining how AI infrastructure is built.

We design and operate AI-native cloud platforms engineered for sovereignty, performance, and scale. Our infrastructure powers GPU-native workloads, multi-tenant control planes, and high-performance AI systems designed for the most demanding environments.

We are not building a generic cloud. We are building purpose-built AI infrastructure - from powered land, to compute, to software .

As we scale our platform and expand our engineering organisation, we are looking for leaders who can build strong teams, uphold high standards, and deliver reliably at pace.

Role Responsibilities
  • Deploy and Manage Kubernetes Clusters, deployed at scale to support AI centric workloads, across both our bare metal clusters and via trusted partner infrastructure
  • Develop Kubernetes Manifests and Operators: Facilitate application deployments and maintain Kubernetes-native services for networking, storage, security, identity and infrastructure management
  • Optimize Linux system configuration including kernel, driver, filesystem and services to support workloads running via our orchestration layer
  • Build and maintain automation scripts and infrastructure as code to support platform lifecycle, as well as simplifying troubleshooting for Incident resolution and provision of tooling for our support organisation
  • Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement.
  • Maintain and enhance Radiant’s observability stack: Prometheus, Grafana, and custom monitoring integrations
  • Operate and support services in 24x7 production environments, including on-call rotation
  • Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations
  • Mentor junior engineers and act as an Operational requirements consultant to other departments
  • Communicate technical decisions clearly to non-technical stakeholders and customers
  • Uphold a culture of: do, document, automate
  • Willingness to cross train with Platform Engineering/Platform SRE to fully support both our infrastructure and platform stacks.
  • Willingness to cross train with HPC Engineering, supported by NVIDIA to enhance our HPC supportability offering
Requirements
  • 5+ Years Proven experience in globally scaled, performance-intensive environments operating to a 24/7 support model in an SRE or equivalent role
  • 3+ years experience in both running, deploying and optimising orchestration platforms with a strong emphasis on Kubernetes
  • Expert-level Linux administration, especially Ubuntu distributions
  • Proficiency in system tuning, disk I/O optimization, and hardware-level performance tweaks
  • Strong networking fundamentals: TCP/IP, DNS, DHCP, VLANs, routing, switching
  • Strong experience with API interrogation
  • Strong experience with infrastructure scripting and automation (Bash, Python, Ansible)
  • Deep understanding of observability principles and tools (Prometheus, Grafana preferred)
  • Strong grasp of ITSM and service operation best practices
  • Excellent communication and mentorship skills
  • Comfortable interfacing with internal stakeholders and external customers
  • Bonus: Knowledge of running AI workloads via orchestration platforms
Bonus Requirements
  • Bachelor or Masters Level degree in Computer Science, Engineering or related field, or equivalent experience.
  • LPIC Certifications
  • ITIL Foundation level qualification or equivalent experience
  • Certified Kubernetes Administrator (CKA)
Qualities we look for:
  • You approach problems with a systems mindset - balancing practical execution with long-term scalability
  • You elevate the team, setting high standards for technical quality and engineering excellence.
  • You hold yourself and others accountable - giving direct feedback and expecting the same
  • You take initiative, owning challenges end-to‑end and proactively driving solutions.
  • You invest in others, mentoring to build both capability and confidence.
Why should you join us?

What sets us apart is our blend of modern technology, competitive benefits, and an open, welcoming work culture that enables our people to thrive.

Here are just some of the great things you can expect from us:

  • 25 days of annual leave
  • A culture that emphasises results over hierarchy, process & ego: we place great emphasis on the quality, ingenuity and creativity of work.
  • Open communication, regular feedback: we value smooth collaboration, direct and actionable feedback, and believe that leading with empathy and a growth mindset makes us better together.
  • Learning Time: we all have dedicated learning time to focus on new skills, projects or interests that lay outside of your day‑to‑day job.
  • Health & Wellbeing: we want everyone to feel healthy and happy, so we offer private medical insurance via Bupa.
  • Cycle to Work Scheme: we're committed to building a sustainable business, so we encourage cycling to work.
  • Gympass subscription to a variety of gyms and wellbeing apps
  • Participation in the company shares program
  • Enhanced parental pay & leave

Diversity, Equality, Inclusion and Belonging

We are an equal opportunity employer and we strive to reduce unconscious bias throughout our hiring process. All applicants will be considered for employment without attention to ethnicity, religion, sexual orientation, gender identity, family or parental status, national origin, veteran, neurodiversity status or disability status. To ensure our recruitment processes provide an equal opportunity for all applicants to succeed, we encourage you to let us know if there are any adjustments that we can make.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Architect
Cluster Architect

Radiant • Greater London

On-site
GBP 120,000 - 190,000
25 days leave
Medical insurance
Cycle to Work
+3
Senior Backend Software Engineer - Backend Systems
Senior Backend Software Engineer - Backend Systems

Radiant • Greater London

On-site
GBP 110,000 - 150,000
25 days leave
Private medical insurance
Cycle to Work
+3
Cloud Infrastructure Support Engineer
Cloud Infrastructure Support Engineer

Radiant • Greater London

On-site
GBP 52,000 - 80,000
25 days annual leave
Private medical insurance (Bupa)
Cycle to Work Scheme
+2
Senior Product Manager - Storage & Networking
Senior Product Manager - Storage & Networking

Jackalope Digital LLC • Greater London

Hybrid
GBP 110,000 - 140,000
25 days annual leave
Health & wellbeing: private medical
Cycle to Work Scheme
+4
Senior Backend Software Engineer - Core Services
Senior Backend Software Engineer - Core Services

Radiant • Greater London

On-site
GBP 50,000 - 95,000
Private medical insurance (Bupa)
Cycle to Work Scheme
Gympass subscription
+3
Senior Product Manager - Storage & Networking
Senior Product Manager - Storage & Networking

Radiant • Greater London

On-site
GBP 90,000 - 130,000
25 days annual leave
Health and wellbeing benefits
Gympass access
+3
Datacentre Operations Engineer
Datacentre Operations Engineer

Radiant • Greater London

On-site
GBP 70,000 - 110,000
On-site in East London
Exposure to NVIDIA GPU AI hardware
Global, multi-discipline engineering
Senior Technical Writer
Senior Technical Writer

Radiant • Greater London

On-site
GBP 65,000 - 90,000
25 days annual leave
Cycle to Work Scheme
Gympass
+3
Senior Technical Writer Radiant London, UK Workplace 4 hours ago
Senior Technical Writer Radiant London, UK Workplace 4 hours ago

Content Creators • Greater London

Hybrid
GBP 65,000 - 95,000
25 days annual leave
Private medical insurance (Bupa)
Cycle to Work Scheme
+3
Senior Manager, Data Center ESG & Sustainability
Senior Manager, Data Center ESG & Sustainability

Radiant • Greater London

On-site
GBP 75,000 - 115,000