HPC & AI Platform Engineer (GPU/Networking)

Carbon3ai Limited.

United Kingdom

Hybrid

GBP 90,000 - 120,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Era4 is seeking a Platform Engineer (HPC & AI) to help shape the new Platform team. The role is customer facing, involves complex troubleshooting, and requires collaboration with vendor engineering teams to ensure seamless AI platform operations.

You will design, deploy and manage large-scale GPU-accelerated clusters using NVIDIA GPUs, Slurm, InfiniBand and high-availability practices, while automating provisioning, monitoring, and security.

Qualifications

  • Experience supporting AI/HPC infrastructure and platforms such as HPE PCAI or similar.
  • System administration with RHEL/CentOS, Ubuntu and kernel tuning.

Responsibilities

  • Designing, deploying, and managing large-scale HPC and GPU-accelerated clusters, including NVIDIA compute environments.
  • Implementing and administering Slurm and related resource-management workflows.
  • Architecting and optimising InfiniBand and Ethernet topologies.
  • Ensuring high availability with failover strategies and proactive maintenance.
  • Automating provisioning, configuration, monitoring, and operational workflows across multi-vendor stacks.
  • Monitoring real-time performance and leading troubleshooting with vendor support.

Skills

HPC infrastructure
NVIDIA GPUs
Slurm
Kernel tuning
Ansible
Kubernetes
CUDA toolkit
InfiniBand networking
RBAC security
Grafana/Kibana
Customer-facing
Vendor coordination

Tools

NVIDIA CUDA toolkit
Kubernetes
Ansible

Job description

Era4 is seeking a Platform Engineer (HPC & AI) to help shape the new Platform team. The role is customer facing, involves complex troubleshooting, and requires collaboration with vendor engineering teams to ensure seamless AI platform operations.

You will design, deploy and manage large-scale GPU-accelerated clusters using NVIDIA GPUs, Slurm, InfiniBand and high-availability practices, while automating provisioning, monitoring, and security.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC & AI Platform Engineer – GPU Clusters
HPC & AI Platform Engineer – GPU Clusters

Era4 • United Kingdom

Hybrid
GBP 95,000 - 130,000
HPC & AI Platform Engineer — GPU Clusters & Automation
HPC & AI Platform Engineer — GPU Clusters & Automation

Era4 • England

On-site
GBP 70,000 - 110,000
Platform Engineer
Platform Engineer

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Data Center Network Engineer for High-Performance AI
Data Center Network Engineer for High-Performance AI

Era4 • England

On-site
GBP 70,000 - 100,000
Senior AI Infrastructure Architect for Scalable GPU and Kubernetes
Senior AI Infrastructure Architect for Scalable GPU and Kubernetes

Carbon3ai Limited. • Greater London

Hybrid
GBP 90,000 - 120,000
Data Centre Network Engineer: AI Infra & Automation
Data Centre Network Engineer: AI Infra & Automation

Carbon3.ai • United Kingdom

On-site
GBP 60,000 - 90,000
Senior AI & HPC Network Architect for Distributed Systems
Senior AI & HPC Network Architect for Distributed Systems

NVIDIA • United Kingdom

On-site
GBP 221,000 - 507,000
AI/HPC Datacenter Solutions Architect
AI/HPC Datacenter Solutions Architect

NVIDIA AI • United Kingdom

On-site
GBP 110,000 - 150,000
AI Data Center Operations Lead — 24/7 HPC/GPU Infra
AI Data Center Operations Lead — 24/7 HPC/GPU Infra

Nscale • Greater London

On-site
GBP 90,000 - 120,000
Cloud Platform Engineer for Scalable AI Workloads
Cloud Platform Engineer for Scalable AI Workloads

Hiverge Ltd • Cambridge

On-site
GBP 70,000 - 115,000