Platform Engineer

Carbon3ai Limited.

United Kingdom

On-site

GBP 90,000 - 120,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Era4 is seeking a Platform Engineer (HPC & AI) to help shape the new Platform team. The role is customer facing, involves complex troubleshooting, and requires collaboration with vendor engineering teams to ensure seamless AI platform operations.

You will design, deploy and manage large-scale GPU-accelerated clusters using NVIDIA GPUs, Slurm, InfiniBand and high-availability practices, while automating provisioning, monitoring, and security.

Qualifications

  • Experience supporting AI/HPC infrastructure and platforms such as HPE PCAI or similar.
  • System administration with RHEL/CentOS, Ubuntu and kernel tuning.

Responsibilities

  • Designing, deploying, and managing large-scale HPC and GPU-accelerated clusters, including NVIDIA compute environments.
  • Implementing and administering Slurm and related resource-management workflows.
  • Architecting and optimising InfiniBand and Ethernet topologies.
  • Ensuring high availability with failover strategies and proactive maintenance.
  • Automating provisioning, configuration, monitoring, and operational workflows across multi-vendor stacks.
  • Monitoring real-time performance and leading troubleshooting with vendor support.

Skills

HPC infrastructure
NVIDIA GPUs
Slurm
Kernel tuning
Ansible
Kubernetes
CUDA toolkit
InfiniBand networking
RBAC security
Grafana/Kibana
Customer-facing
Vendor coordination

Tools

NVIDIA CUDA toolkit
Kubernetes
Ansible

Job description

Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data‑centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations

Role Summary:

We are looking for Platform Engineer (HPC & AI) who can assist in shaping our new Platform team, this role will be customer facing, involve technical troubleshooting, and collaboration with vendor engineering teams to ensure seamless AI platform operations.

Responsibilities:
  • Designing, deploying, and managing large‑scale HPC and GPU‑accelerated clusters, including NVIDIA based compute environments.
  • Implementing and administering HPC scheduling and resource‑management systems (e.g., Slurm), including GPU partitioning, workload scheduling, and capacity planning.
  • Architecting and optimising InfiniBand and Ethernet network topologies.
  • Ensuring high availability and resilience through failover strategies, planned maintenance coordination, and proactive risk mitigation.
  • Automating provisioning, configuration, monitoring, and operational workflows across multi‑vendor HPC hardware and software stacks.
  • Monitoring real‑time performance and leading troubleshooting efforts across compute, storage, interconnect, drivers, and node failures, engaging vendor support for critical issues.
  • Incident response: node failure management, network issues, driver issues, troubleshooting common issues and then working with vendor support to resolve any critical issues.
  • Security and access control: Manage user permissions, RBAC, security hardening, data protection.
Required Skills & Experience:
  • Experience supporting HPE PCAI or other AI/HPC infrastructure and platforms.
  • System administration experience with OS's like RHEL/CentOS, Ubuntu, tuning Linux kernel.
  • Proficiency with Ansible, Nvidia and CUDA toolkits, Kubernetes and container orchestration.
  • Understanding of automation, monitoring and security with GPU as a service.
  • Extensive experience in system engineering, platform operations or SRE.
  • Experience with GPU resource allocation (across instances, GPUs count and time).
  • Advanced networking skills with High performance networking, troubleshooting and fine tuning.
  • Familiarity with cloud-based platforms, APIs, and distributed systems.
  • Understanding of AI/ML concepts and tooling (model training, inference, data pipelines basics).
  • Experience with monitoring/logging tools (e.g., Grafana, Kibana, Splunk).
  • Excellent communication skills to interface with both customers and internal / vendor teams.
  • Good understanding of tools requirements for ML engineers and data scientists, and how to optimise the experience.
Why Join Era4:

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next‑generation company operates at scale.

Diversity & Inclusion :

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Note:

We appreciate this is a relatively new skill set and we are open to candidates who may not tick all the boxes but are willing to learn and develop their skillset.

Technology

United Kingdom: (Occasional office visit required)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Solutions Architect - AI Infrastructure
Solutions Architect - AI Infrastructure

Carbon3ai Limited. • Greater London

Hybrid
GBP 90,000 - 120,000
HPC & AI Platform Engineer – GPU Clusters
HPC & AI Platform Engineer – GPU Clusters

Era4 • United Kingdom

Hybrid
GBP 95,000 - 130,000
HPC & AI Platform Engineer — GPU Clusters & Automation
HPC & AI Platform Engineer — GPU Clusters & Automation

Era4 • England

On-site
GBP 70,000 - 110,000
HPC & AI Platform Engineer (GPU/Networking)
HPC & AI Platform Engineer (GPU/Networking)

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Technical Lead, HPC & AI Infrastructure
Technical Lead, HPC & AI Infrastructure

Era4 • England

On-site
GBP 90,000 - 150,000
NOC Engineer (24x7)
NOC Engineer (24x7)

Era4 • West of England

On-site
GBP 32,000 - 45,000
InfoSec Analyst
InfoSec Analyst

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 55,000 - 75,000
Senior AI Infrastructure Architect for Scalable GPU and Kubernetes
Senior AI Infrastructure Architect for Scalable GPU and Kubernetes

Carbon3ai Limited. • Greater London

Hybrid
GBP 90,000 - 120,000
Network Engineer
Network Engineer

asobbi • United Kingdom

Remote
GBP 53,000 - 69,000
Highly competitive package with equity
Dynamic progression plan
Human-first flexibility
Site Engineer
Site Engineer

Era4 • Chesterfield

On-site
GBP 40,000 - 60,000