East London HPC Data Centre Operations Engineer

Radiant

Greater London

On-site

GBP 70,000 - 110,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

On-site in East London
Exposure to NVIDIA GPU AI hardware
Global, multi-discipline engineering

Job summary

Radiant in East London is building an AI infrastructure site and seeks a Datacentre Operations Engineer to execute on-the-ground hardware tasks and ensure uptime for high-density compute across East London facilities.

You will work with SRE, Network Engineering, and Datacentre Strategy teams to deploy, maintain, and troubleshoot GPU/CPU platforms, cooling, and power, driving operational excellence while scaling across the region.

Qualifications

  • Degree in Computer Science or Electrical Engineering, or 5+ years of directly relevant data centre operations experience.
  • 3+ years of experience in data centre operations, HPC, or related roles.
  • Strong communication skills in English.
  • Proven hands-on experience with HPC/NVIDIA GPU platforms or equivalent high-density compute systems.

Responsibilities

  • Diagnose and resolve hardware and network issues to maximise uptime; execute structured fault isolation methodologies to drive rapid resolution
  • Respond to critical hardware alerts via our monitoring and observability platform; contribute to ongoing service improvement to improve monitoring capability and alert quality
  • Deploy and maintain HPC and AI hardware for uninterrupted operations, including hardware troubleshooting, firmware updates, and component replacement
  • Execute break/fix procedures for advanced hardware platforms, including GPU module exchange, component-level fault isolation, and firmware-level diagnostics
  • Execute or support break/fix operations on ultra-high-density compute systems including NVIDIA B300-class chassis or equivalent platforms, including GPU/fan module exchange, chassis-level fault isolation, and busbar connection/disconnection—under the direction of the Lead where qualification is in progress
  • Operate, monitor, and maintain high-density air cooling infrastructure in conjunction with our datacentre partner, including CRAC/CRAH units, in-row and containment cooling, and associated airflow management systems
  • Facilitate in conjunction with our datacentre partner, routine and corrective maintenance on air cooling systems: monitoring supply/return temperatures and airflow rates, maintaining hot-aisle/cold-aisle containment and blanking, and performing scheduled filter and component inspections
  • Follow and contribute to SOPs for safe working around high-density, air-cooled compute platforms
  • Monitor thermal performance and raise anomalies before they elevate into incidents
  • Contribute to site-level capacity management operations, maintaining accurate records of power, space, and cooling utilisation
  • Support capacity planning activities by providing accurate as-built data and flagging infrastructure changes to the Lead and relevant teams
  • Manage on- ground assets from point of purchase and delivery through lifecycle management and disposal, owning asset management within Radiant's CMDB system
  • Handle RMAs and support requests within Radiant's Service Level Objectives (SLOs) to meet customer contract SLAs
  • Contribute to ongoing maintenance, fostering compliance and leveraging strong vendor partnerships
  • Operate cooling, power distribution (including busbar and PDU infrastructure), and other critical data centre technologies to maintain high operational standards
  • Develop and maintain datacentre/hardware management SOPs, ensuring continual alignment with Radiant's governance and compliance requirements
  • Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement
  • Operate and support services 24x7x365 for production environments as part of a structured on-site shift rotation—working 12-hour shifts across a 4-team pattern with two engineers on shift at all times—to meet tight customer SLAs
  • Prioritise and triage incident and smart-hands workload ensuring tight SLA coverage is maintained with a small on the ground team
  • Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations
  • Communicate technical decisions clearly to stakeholders and customers
  • Champion a culture of: do, document, automate
  • Willing to cross train and upskill in Infrastructure/Platform SRE practices
  • Willing to travel across EMEA to support future datacentre onboarding and train in new technologies

Job description

Radiant in East London is building an AI infrastructure site and seeks a Datacentre Operations Engineer to execute on-the-ground hardware tasks and ensure uptime for high-density compute across East London facilities.

You will work with SRE, Network Engineering, and Datacentre Strategy teams to deploy, maintain, and troubleshoot GPU/CPU platforms, cooling, and power, driving operational excellence while scaling across the region.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Datacentre Operations Engineer
Datacentre Operations Engineer

Radiant • Greater London

On-site
GBP 70,000 - 110,000
On-site in East London
Exposure to NVIDIA GPU AI hardware
Global, multi-discipline engineering
24/7 Cloud Infra Support Engineer for AI & HPC
24/7 Cloud Infra Support Engineer for AI & HPC

Radiant • Greater London

On-site
GBP 52,000 - 80,000
25 days annual leave
Private medical insurance (Bupa)
Cycle to Work Scheme
+2
AI Data Center Engineer: GPU HPC & Networking
AI Data Center Engineer: GPU HPC & Networking

Hamilton Barnes Associates Limited • United Kingdom

On-site
GBP 45,000 - 75,000
Data center Technician
Data center Technician

WeEngage Group • Wallsend

On-site
GBP 24,000 - 32,000
Competitive salary
Healthcare
Training and career development
+2
AI Data Center Operations Lead - Equity & Growth
AI Data Center Operations Lead - Equity & Growth

Hamilton Barnes Associates Limited • United Kingdom

On-site
GBP 67,500 - 82,500
Cash + equity compensation
Healthcare, wellbeing and lunch benef
Profitable and rapidly growing organ
+3
Data Centre Operations Engineer - GPU/Linux Infra
Data Centre Operations Engineer - GPU/Linux Infra

Easy Compute • Bury

On-site
GBP 35,000 - 40,000
Head of AI Data Centre Deployment & Operations
Head of AI Data Centre Deployment & Operations

asobbi • United Kingdom

Hybrid
GBP 150,000 - 200,000
Competitive compensation
Potential equity
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability

Radiant • England

On-site
GBP 70,000 - 90,000
Exposure to industry-leading GPU and AI infrastructure
Collaborative, inclusive, and supportive engineering culture
Real ownership and influence over operational excellence
Data Centre Technical Manager: Operations & Eng Lead
Data Centre Technical Manager: Operations & Eng Lead

PRS • Greater London

On-site
GBP 80,000 - 85,000
Car allowance
Annual bonus
Management benefits
Platform Engineer
Platform Engineer

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000