Platform Engineer (7351A69)

Referment

Greater London

On-site

GBP 90,000 - 120,000

Full time

13 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Referment is building an on‑prem Kubernetes platform to underpin its global, low‑latency trading workloads. You will design and operate the VM fabric (KVM/libvirt) and the full Kubernetes lifecycle—from bootstrapping and upgrades to storage integration and disaster recovery.

You will own multi‑tenancy patterns, implement GitOps with ArgoCD or Flux, and harden environments against CIS benchmarks while collaborating with datacentre and network teams to ensure secure, resilient operation.

Qualifications

  • Solid production experience running Kubernetes in a self‑managed on‑premises context.
  • Strong Linux systems administration background: networking, storage, kernel troubleshooting.
  • Experience with infrastructure‑as‑code (Terraform, Ansible, Packer) and GitOps workflows.
  • Familiar with observability stacks (Prometheus/Grafana, ELK) and CIS benchmarks.

Responsibilities

  • Design, build and operate the on‑prem Kubernetes platform beneath application workloads.
  • Manage VM fabric (KVM/libvirt), hypervisor architecture and network/storage backing.
  • Own CI/CD pipelines integration, install/upgrade of clusters and disaster recovery.
  • Coordinate with data centre and network teams for physical provisioning and on‑call incidents.

Skills

Kubernetes
Linux administration
Networking
Storage
Terraform
Ansible
Packer
GitOps
Prometheus/Grafana

Tools

kubeadm
Cluster API
Kubespray

Job description

Referment is working with a global options market maker whose low‑latency trading platform runs on self‑managed, on‑premises infrastructure rather than public cloud. The firm trades across global derivatives markets, builds its platform on open‑source tooling where it fits, and values a flat, collaborative engineering culture.

The Role

You will help design, build and operate the on‑premises Kubernetes platform that sits beneath the firm's application workloads — from the hypervisor and VM provisioning up through Kubernetes cluster lifecycle, networking, storage and the golden-path tooling application teams use to ship software. The Linux VM fabric (KVM/libvirt‑based) is still being designed and built out, so you will have real influence over its architecture, not just its day‑to‑day operation, and you will keep the fabric and the clusters running on top of it healthy, secure and performant without the safety net of a hyperscaler's managed control plane.

Your work will span the full stack: designing the VM fabric’s hypervisor architecture, host networking topology and storage backing; designing and maintaining distributed shared storage such as Ceph; building VM templating and golden‑image pipelines; automating the VM lifecycle end to end via infrastructure‑as‑code; and managing capacity planning, oversubscription strategy and headroom for failure or maintenance. You will own virtual networking within the fabric and its clean handoff into the Kubernetes CNI layer; design, build and maintain the full lifecycle of on‑prem Kubernetes clusters (bootstrapping, version upgrades, node scaling, decommissioning) using tooling such as kubeadm, Cluster API or Kubespray; manage the control plane end to end, including etcd operations, backup/restore, performance tuning and disaster recovery; configure CNI, network policy enforcement and on‑prem load balancing; stand up ingress and internal DNS; and own persistent storage integration via CSI drivers. You will define and enforce multi‑tenancy patterns across both layers, build GitOps‑based delivery with ArgoCD or Flux, harden hosts and clusters against CIS benchmarks, manage secrets, build observability across the full stack from hypervisor health to cluster metrics and logs, plan and execute upgrades with minimal workload disruption, troubleshoot incidents from a misbehaving pod down through kubelet, the container runtime and CNI into the underlying VM and hypervisor, share an on‑call rotation for platform‑level incidents, partner with application teams on developer experience, and coordinate with datacentre and network teams on physical host provisioning.

What We're Looking For
  • Solid production experience running Kubernetes in a self‑managed, on‑premises context (not just EKS/GKE/AKS) — you understand what breaks when there's no managed control plane, and you've operated etcd and the control plane yourself
  • Hands‑on experience designing and operating a Linux KVM/libvirt‑based VM fabric as the foundation for Kubernetes, ideally from a relatively early stage rather than just inheriting a mature environment
  • Strong Linux systems administration background: networking, storage, service management, kernel tuning and troubleshooting under pressure
  • Practical, in‑depth knowledge of Kubernetes networking (CNI internals; service meshes a plus) and storage (CSI drivers, distributed storage systems such as Ceph/Longhorn)
  • Experience with infrastructure‑as‑code and configuration management (Terraform, Ansible, Packer), GitOps workflows and CI/CD pipelines
  • Comfortable with observability stacks (Prometheus/Grafana, ELK/Loki) and using them to diagnose infrastructure issues without cloud‑native tooling
  • Security‑conscious: familiar with hardening standards, RBAC, network segmentation and secrets management
  • Strong troubleshooting instincts across the full stack — hypervisor, OS, network, container runtime, Kubernetes control plane — and good written and verbal communication with distributed or hybrid teams
Relevant Desirable Experience
  • Bare‑metal Kubernetes provisioning; a background in a regulated or air‑gapped/restricted‑network environment
  • Contributions to open‑source infrastructure tooling
  • Using AI tooling to accelerate development; experience running AI infrastructure

This could suit an infrastructure or platform engineer who likes understanding systems all the way down — from hypervisor to pod — and wants real ownership with no black‑box managed services standing between them and root cause.

#Referment

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform Engineer
Senior Platform Engineer

Selby Jennings • Greater London

On-site
GBP 90,000 - 130,000
Platform Engineer
Platform Engineer

Avanti • United Kingdom

Remote
GBP 70,000 - 95,000
Remote-first working
25 days' annual leave
Company pension
+3
Kubernetes Platform Engineer
Kubernetes Platform Engineer

G-Research • Greater London

On-site
GBP 90,000 - 120,000
Highly competitive compensation
Annual discretionary bonus
Lunch provided
+5
Platform Engineer – Graduate Considered
Platform Engineer – Graduate Considered

Redtech Recruitment Ltd. • Cambourne

Hybrid
GBP 42,000 - 68,000
Platform Reliability Engineer — IaC, Kubernetes & GCP
Platform Reliability Engineer — IaC, Kubernetes & GCP

Selby Jennings • Greater London

On-site
GBP 90,000 - 130,000
Kubernetes Manager
Kubernetes Manager

Intelix.AI • Greater London

On-site
GBP 99,000 - 121,000
£22k bonus
15% Pension
30 Days Holiday
+1
Senior Platform Engineer
Senior Platform Engineer

TrueNorth® • Greater London

Hybrid
GBP 100,000 - 130,000
Share options
Onsite allowance
Platform Engineer – Cloud & Kubernetes
Platform Engineer – Cloud & Kubernetes

Talenzon group • Greater London

On-site
GBP 60,000 - 85,000
Platform Engineer
Platform Engineer

Searchability® • Sheffield

On-site
GBP 65,000 - 79,000
Platform Engineer
Platform Engineer

Systematica Group • Greater London

On-site
GBP 90,000 - 130,000