Platform Engineer - HPC & AI Infra Architect

Carbon3.ai

United Kingdom

On-site

GBP 85,000 - 120,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Carbon3.ai is seeking a Platform Engineer (HPC & AI) in the United Kingdom to shape our new Platform team. The role involves customer-facing collaboration, technical troubleshooting, and coordination with vendor engineering teams to ensure seamless AI platform operations.

You will design, deploy, and manage large‑scale HPC clusters, implement Slurm governance, optimize network topologies, and automate workflows across multi‑vendor stacks.

Qualifications

  • Experience supporting AI/HPC infrastructure and platforms.
  • System administration experience on Linux (RHEL/CentOS/Ubuntu) and kernel tuning.
  • Proficiency with Ansible and GPU toolkits (NVIDIA CUDA).
  • Understanding automation, monitoring, and security for GPU-as-a-service.
  • Extensive experience in platform operations or SRE.
  • Experience with GPU resource allocation across instances and time.
  • Advanced networking skills for high‑performance networks.
  • Familiarity with cloud APIs and distributed systems.
  • Knowledge of AI/ML concepts, pipelines, and tooling.
  • Experience with monitoring/logging tools (Grafana, Kibana, Splunk).

Responsibilities

  • Design, deploy, and manage large-scale HPC and GPU-accelerated clusters.
  • Administer HPC scheduling systems (e.g., Slurm) including GPU partitioning.
  • Architect InfiniBand and Ethernet network topologies.
  • Ensure high availability and proactive risk mitigation.
  • Automate provisioning, configuration, and monitoring across multi-vendor stacks.
  • Monitor real-time performance and troubleshoot compute/storage issues.
  • Respond to incidents including node, network, and driver issues.
  • Manage access control, RBAC, and security hardening.

Skills

HPC & AI platform support
Linux system administration
Ansible
NVIDIA CUDA toolkit
Kubernetes
Networking (HDI)
Monitoring & observability
Security hardening
Vendor collaboration
GPU resource management

Tools

Slurm
InfiniBand
Grafana
Kubernetes

Job description

Carbon3.ai is seeking a Platform Engineer (HPC & AI) in the United Kingdom to shape our new Platform team. The role involves customer-facing collaboration, technical troubleshooting, and coordination with vendor engineering teams to ensure seamless AI platform operations.

You will design, deploy, and manage large‑scale HPC clusters, implement Slurm governance, optimize network topologies, and automate workflows across multi‑vendor stacks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC & AI Platform Engineer – GPU Clusters
HPC & AI Platform Engineer – GPU Clusters

Era4 • United Kingdom

Hybrid
GBP 95,000 - 130,000
HPC & AI Platform Engineer (GPU/Networking)
HPC & AI Platform Engineer (GPU/Networking)

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Platform Engineer
Platform Engineer

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Platform Engineer: Cloud Infra & AI Deployment
Platform Engineer: Cloud Infra & AI Deployment

Jackalope Digital LLC • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Cloud Platform Engineer for AI Infrastructure
Senior Cloud Platform Engineer for AI Infrastructure

Graphcore • West of England

On-site
GBP 70,000 - 120,000
Flexible working
Annual leave
Private medical insurance
+6
HPC & AI Platform Engineer — GPU Clusters & Automation
HPC & AI Platform Engineer — GPU Clusters & Automation

Era4 • England

On-site
GBP 70,000 - 110,000
Staff Cloud Platform Engineer for AI Compute
Staff Cloud Platform Engineer for AI Compute

EngineersOfAI • Bristol

On-site
GBP 60,000 - 80,000
Staff Cloud Platform Engineer — AI Infra & OpenStack
Staff Cloud Platform Engineer — AI Infra & OpenStack

Graphcore • West of England

Hybrid
GBP 85,000 - 125,000
Flexible working
Annual leave policy
Private medical insurance
+3
MLOps & Infrastructure Engineer for AI HPC
MLOps & Infrastructure Engineer for AI HPC

EngineersOfAI • United Kingdom

On-site
GBP 40,000 - 70,000
AI Platform Engineer: Scalable Model Serving & Infra
AI Platform Engineer: Scalable Model Serving & Infra

Barlowe LLP • Greater London

On-site
GBP 70,000 - 90,000
Highly competitive compensation
30 days annual leave
9% company pension contributions
+4