An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Referment is building an on‑prem Kubernetes platform to underpin its global, low‑latency trading workloads. You will design and operate the VM fabric (KVM/libvirt) and the full Kubernetes lifecycle—from bootstrapping and upgrades to storage integration and disaster recovery.
You will own multi‑tenancy patterns, implement GitOps with ArgoCD or Flux, and harden environments against CIS benchmarks while collaborating with datacentre and network teams to ensure secure, resilient operation.
Referment is working with a global options market maker whose low‑latency trading platform runs on self‑managed, on‑premises infrastructure rather than public cloud. The firm trades across global derivatives markets, builds its platform on open‑source tooling where it fits, and values a flat, collaborative engineering culture.
You will help design, build and operate the on‑premises Kubernetes platform that sits beneath the firm's application workloads — from the hypervisor and VM provisioning up through Kubernetes cluster lifecycle, networking, storage and the golden-path tooling application teams use to ship software. The Linux VM fabric (KVM/libvirt‑based) is still being designed and built out, so you will have real influence over its architecture, not just its day‑to‑day operation, and you will keep the fabric and the clusters running on top of it healthy, secure and performant without the safety net of a hyperscaler's managed control plane.
Your work will span the full stack: designing the VM fabric’s hypervisor architecture, host networking topology and storage backing; designing and maintaining distributed shared storage such as Ceph; building VM templating and golden‑image pipelines; automating the VM lifecycle end to end via infrastructure‑as‑code; and managing capacity planning, oversubscription strategy and headroom for failure or maintenance. You will own virtual networking within the fabric and its clean handoff into the Kubernetes CNI layer; design, build and maintain the full lifecycle of on‑prem Kubernetes clusters (bootstrapping, version upgrades, node scaling, decommissioning) using tooling such as kubeadm, Cluster API or Kubespray; manage the control plane end to end, including etcd operations, backup/restore, performance tuning and disaster recovery; configure CNI, network policy enforcement and on‑prem load balancing; stand up ingress and internal DNS; and own persistent storage integration via CSI drivers. You will define and enforce multi‑tenancy patterns across both layers, build GitOps‑based delivery with ArgoCD or Flux, harden hosts and clusters against CIS benchmarks, manage secrets, build observability across the full stack from hypervisor health to cluster metrics and logs, plan and execute upgrades with minimal workload disruption, troubleshoot incidents from a misbehaving pod down through kubelet, the container runtime and CNI into the underlying VM and hypervisor, share an on‑call rotation for platform‑level incidents, partner with application teams on developer experience, and coordinate with datacentre and network teams on physical host provisioning.
This could suit an infrastructure or platform engineer who likes understanding systems all the way down — from hypervisor to pod — and wants real ownership with no black‑box managed services standing between them and root cause.
#Referment