Senior HPC Systems Engineer - GPU & Bare-Metal Kubernetes

remotestar-team

Cambourne

Hybrid

GBP 90,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Indefinite contract
Equal pay guaranteed
Variable performance bonus
Signing bonus
Relocation package
Private health insurance
Educational budget
Hybrid opportunity
Flexible working hours
High-paced cutting-edge tech

Job summary

Unknown company is seeking a senior systems engineer to design and build the control plane for bare-metal AI infrastructure. You will architect GPU scheduling across massive clusters, optimize networking for low latency, and extend Kubernetes with Operators and CRDs to expose hardware realities to data scientists.

Deep kernel and device driver work is required. You will debug complex issues from PCIe to NCCL timeouts, and define standards for drivers, firmware, and OS optimizations to maximize

Qualifications

  • 10+ years of software engineering experience with Python.
  • Deep Kubernetes knowledge: CRDs, Operators, API server architecture.
  • Hands-on experience managing NVIDIA GPU clusters, drivers, CUDA, and NVIDIA Container Toolkit.
  • Deep Linux kernel, cgroups, namespaces, and performance tuning.
  • Terraform/Ansible for provisioning physical hardware.
  • Experience debugging distributed systems across code, network, or silicon.

Responsibilities

  • Build the Control Plane to automate bare-metal AI infrastructure lifecycles.
  • Architect scheduling for large-scale GPU clusters with efficient bin-packing and gang scheduling.
  • Tune the software-defined networking layer for low-latency interconnects (InfiniBand/RDMA).
  • Develop Kubernetes Extensions via Operators and CRDs for hardware abstractions.
  • Debug PCIe, NCCL timeouts, and kernel panics on bare-metal nodes.
  • Define the Golden Image for AI workloads including drivers and firmware.

Skills

Python
Kubernetes
Linux Internals
GPU Clusters
Infrastructure as Code
Troubleshooting

Tools

NVIDIA Container Toolkit
Cluster API (CAPI)
Metal3
Tinkerbell
Canonical MaaS
OpenStack Ironic

Job description

Unknown company is seeking a senior systems engineer to design and build the control plane for bare-metal AI infrastructure. You will architect GPU scheduling across massive clusters, optimize networking for low latency, and extend Kubernetes with Operators and CRDs to expose hardware realities to data scientists.

Deep kernel and device driver work is required. You will debug complex issues from PCIe to NCCL timeouts, and define standards for drivers, firmware, and OS optimizations to maximize

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC

NVIDIA • United Kingdom

On-site
GBP 110,000 - 170,000
Senior GPU & AI Infrastructure Architect
Senior GPU & AI Infrastructure Architect

Referment • Greater London

On-site
GBP 90,000 - 120,000
Senior GPU HPC Engineer: InfiniBand & KVM Optimization
Senior GPU HPC Engineer: InfiniBand & KVM Optimization

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior Data Center Engineer — GPU & Linux Networking
Senior Data Center Engineer — GPU & Linux Networking

Pursuu • Manchester

Hybrid
GBP 40,000 - 70,000
Company events
Company pension
Free parking
+2
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

remotestar-team • Cambourne

Hybrid
GBP 90,000 - 150,000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+7
Lead GPU Infrastructure Architect for Scalable AI Clusters
Lead GPU Infrastructure Architect for Scalable AI Clusters

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
Senior GPU Solutions Architect – Enterprise HPC UK
Senior GPU Solutions Architect – Enterprise HPC UK

Referment • Greater London

On-site
GBP 90,000 - 150,000
Platform Engineer — GPU HPC & Bare-Metal Clusters
Platform Engineer — GPU HPC & Bare-Metal Clusters

CATCHES • United Kingdom

Remote
GBP 50,000 - 70,000
Senior GPU & AI Infra Architect — Remote, 4-Day Week
Senior GPU & AI Infra Architect — Remote, 4-Day Week

Civo Ltd • United Kingdom

Hybrid
GBP 110,000 - 170,000
4-day week
Uncapped holidays
Remote work environment
HPC IT Lead Engineer for GPU/CPU Simulations
HPC IT Lead Engineer for GPU/CPU Simulations

Linuxconfig • Greater London

Hybrid
GBP 90,000 - 130,000
Equity options
10% employer pension contribution
Free office lunches
+10