HPC & AI Platform Engineer — GPU Clusters & Automation

Era4

England

On-site

GBP 70,000 - 110,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Era4 develops and operates AI infrastructure across the UK, turning brownfield sites into modern data-centre facilities for healthcare, research, finance, enterprise and public sector organisations.

The Platform Engineer role shapes Era4's new Platform team, is customer-facing, and involves troubleshooting and cross-vendor collaboration to ensure seamless AI platform operations at scale.

Qualifications

  • Experience supporting AI/HPC infrastructure and platforms.
  • Linux system administration with RHEL/CentOS and Ubuntu.
  • Proficiency in NVIDIA CUDA toolkits, Kubernetes, and Ansible.
  • Familiarity with high-performance networking and security hardening.
  • Ability to troubleshoot, automate, and collaborate with vendor teams.

Responsibilities

  • Design, deploy, and manage large-scale HPC and GPU clusters.
  • Implement HPC scheduling and resource management (Slurm).
  • Architect InfiniBand/Ethernet topologies and optimise networks.
  • Ensure high availability and proactive risk mitigation.
  • Automate provisioning, monitoring, and workflows across multi-vendor stacks.
  • Lead real-time performance monitoring and fault diagnosis with vendor support.
  • Manage security, RBAC, and data protection; interface with customers.

Skills

AI/HPC platform experience
System administration
Linux admin
NVIDIA CUDA toolkits
Kubernetes & containers
Automation with Ansible
HPC networking
Cloud APIs
Monitoring Grafana
Security RBAC
Vendor collaboration

Tools

Ansible
NVIDIA CUDA toolkit
Kubernetes
Grafana
Kibana
Splunk

Job description

Era4 develops and operates AI infrastructure across the UK, turning brownfield sites into modern data-centre facilities for healthcare, research, finance, enterprise and public sector organisations.

The Platform Engineer role shapes Era4's new Platform team, is customer-facing, and involves troubleshooting and cross-vendor collaboration to ensure seamless AI platform operations at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC & AI Platform Engineer – GPU Clusters
HPC & AI Platform Engineer – GPU Clusters

Era4 • United Kingdom

Hybrid
GBP 95,000 - 130,000
HPC & AI Platform Engineer (GPU/Networking)
HPC & AI Platform Engineer (GPU/Networking)

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Platform Engineer
Platform Engineer

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Senior AI Infrastructure Architect for Scalable GPU and Kubernetes
Senior AI Infrastructure Architect for Scalable GPU and Kubernetes

Carbon3ai Limited. • Greater London

Hybrid
GBP 90,000 - 120,000
Data Center Network Engineer for High-Performance AI
Data Center Network Engineer for High-Performance AI

Era4 • England

On-site
GBP 70,000 - 100,000
Solutions Architect - AI Infrastructure
Solutions Architect - AI Infrastructure

Carbon3ai Limited. • Greater London

Hybrid
GBP 90,000 - 120,000
Data Centre Network Engineer: AI Infra & Automation
Data Centre Network Engineer: AI Infra & Automation

Carbon3.ai • United Kingdom

On-site
GBP 60,000 - 90,000
Security Engineer: Cloud & Kubernetes for AI Infra
Security Engineer: Cloud & Kubernetes for AI Infra

Era4 • England

On-site
GBP 70,000 - 90,000
Platform Engineer — GPU HPC & Bare-Metal Clusters
Platform Engineer — GPU HPC & Bare-Metal Clusters

CATCHES • United Kingdom

Remote
GBP 50,000 - 70,000
Platform Engineer - AI Infra & Large-Scale GPU Systems
Platform Engineer - AI Infra & Large-Scale GPU Systems

Ineffable • Greater London

On-site
GBP 90,000 - 110,000