HPC & AI Platform Engineer – GPU Clusters

Era4

United Kingdom

Hybrid

GBP 95,000 - 130,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Era4 is seeking a Platform Engineer (HPC & AI) in the United Kingdom to shape a new Platform team. The role is customer facing, involves troubleshooting, and collaboration with vendor engineering teams to ensure seamless AI platform operations.

You will design, deploy, and manage large-scale HPC and GPU-accelerated clusters (NVIDIA). You will implement Slurm-based scheduling, optimize InfiniBand/Ethernet networks, and lead automation, monitoring, and incident response across multi-vendor stacks

Qualifications

  • Experience supporting AI/HPC infrastructure and platforms.
  • System administration experience with Linux (RHEL/CentOS, Ubuntu).
  • Proficiency with Ansible, Nvidia CUDA toolkits, Kubernetes; automation and security focus.

Responsibilities

  • Design, deploy, and manage large-scale HPC and GPU-accelerated clusters (NVIDIA).
  • Implement and administer HPC scheduling with GPU partitioning (e.g., Slurm).
  • Architect and optimise InfiniBand and Ethernet network topologies.
  • Ensure high availability and resilience with failover and maintenance planning.
  • Automate provisioning, configuration, monitoring, and workflows across multi-vendor stacks.
  • Monitor real-time performance and lead troubleshooting with vendor support.
  • Incident response: node/network/driver issues and escalation with vendors.
  • Security and access control: manage permissions, RBAC, data protection.

Skills

HPC/AI infra
Linux admin
Ansible
NVIDIA CUDA
Kubernetes
GPU scheduling
SLURM
High perf networking
RBAC security

Tools

NVIDIA CUDA toolkit
Kubernetes
Ansible

Job description

Era4 is seeking a Platform Engineer (HPC & AI) in the United Kingdom to shape a new Platform team. The role is customer facing, involves troubleshooting, and collaboration with vendor engineering teams to ensure seamless AI platform operations.

You will design, deploy, and manage large-scale HPC and GPU-accelerated clusters (NVIDIA). You will implement Slurm-based scheduling, optimize InfiniBand/Ethernet networks, and lead automation, monitoring, and incident response across multi-vendor stacks

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC & AI Platform Engineer (GPU/Networking)
HPC & AI Platform Engineer (GPU/Networking)

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
HPC & AI Platform Engineer — GPU Clusters & Automation
HPC & AI Platform Engineer — GPU Clusters & Automation

Era4 • England

On-site
GBP 70,000 - 110,000
Platform Engineer
Platform Engineer

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Senior AI Infrastructure Architect for Scalable GPU and Kubernetes
Senior AI Infrastructure Architect for Scalable GPU and Kubernetes

Carbon3ai Limited. • Greater London

Hybrid
GBP 90,000 - 120,000
Data Center Network Engineer for High-Performance AI
Data Center Network Engineer for High-Performance AI

Era4 • England

On-site
GBP 70,000 - 100,000
Senior GPU Systems Engineer - Clusters & Platform Infra
Senior GPU Systems Engineer - Clusters & Platform Infra

Radley James • Greater London

On-site
GBP 20,000 - 40,000
Platform Engineer — GPU HPC & Bare-Metal Clusters
Platform Engineer — GPU HPC & Bare-Metal Clusters

CATCHES • United Kingdom

Remote
GBP 50,000 - 70,000
Senior GPU HPC Engineer: InfiniBand & KVM Optimization
Senior GPU HPC Engineer: InfiniBand & KVM Optimization

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Platform Engineer - AI Infra & Large-Scale GPU Systems
Platform Engineer - AI Infra & Large-Scale GPU Systems

Ineffable • Greater London

On-site
GBP 90,000 - 110,000
Senior GPU & AI Infra Architect — Remote, 4-Day Week
Senior GPU & AI Infra Architect — Remote, 4-Day Week

Civo Ltd • United Kingdom

Hybrid
GBP 110,000 - 170,000
4-day week
Uncapped holidays
Remote work environment