Senior Infrastructure Engineer — GPU Compute

Circle B

Hoofddorp

On-site

EUR 90,000 - 130,000

Full time

20 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Company gym
Modern office in Hoofddorp
Informal working culture

Job summary

Circle B builds a sovereign EU GPU cloud and modern AI infrastructure. You will own bare-metal Kubernetes, manage the NVIDIA GPU stack, and secure host-side networking to deliver GPU-ready clusters for tenants.

Join a hardware-focused team that values automation, security, and GitOps-driven operations while working from a modern Hoofddorp office.

Qualifications

  • Minimum 5+ years in infrastructure or SRE.
  • Minimum 3+ years running Kubernetes in production with lifecycle ownership.
  • Strong Linux fundamentals (systemd, kernel modules, NUMA, cgroups, NVMe/LVM).
  • Hands-on server provisioning via IPMI/Redfish/BMC tooling.
  • Experience with NVIDIA GPU stack on Kubernetes and GPU health telemetry.
  • IaC and GitOps in production (Helm, Kustomize, Argo CD or Flux).
  • Scripting in Python and/or Go; self-directed problem solving.

Responsibilities

  • Manage the HA management cluster on the provisioning hardware.
  • Oversee automated bare-metal provisioning for GPU fleet (BMC/Redfish, PXE).
  • Operate NVIDIA GPU stack on Kubernetes (driver, container toolkit, device plugin).
  • Handle host GPU fabric networking (NVLink/NVSwitch, RDMA NICs).
  • Manage secrets and certificates (Vault/OpenBao, cert-manager).
  • Implement GitOps-driven provisioning; ensure repeatability and auditability.
  • Oversee secure tenant decommissioning and GPU memory zeroing.

Skills

Linux fundamentals
Kubernetes in production
IPMI/Redfish/BMC tooling
Infrastructure as code
GitOps (Argo CD/Flux)
Python/Go scripting
Self-directed problem solver

Tools

NVIDIA GPU Operator
Vault/OpenBao/Cert-manager
Ansible
Docker/Container tooling

Job description

Circle B builds sustainable IT infrastructure for the AI and cloud era. For over a decade— we have designed and deployed datacenter, edge, and AI/HPC systems on Open Compute Project (OCP) hardware. We are independent, vendor-neutral, and ISO 9001 / 14001 / 27001 certified, with deployments across multiple countries.

Our newest initiative is a sovereign EU GPU cloud — operating under full Dutch/EU jurisdiction and beyond the reach of the US CLOUD Act, for regulated European organizations that cannot compromise on where their data lives.

The Role

You look after the servers and the clusters that run on them: bringing machines up from their BMC, provisioning them through automated tooling, and handing over GPU-ready Kubernetes clusters for tenants to use.

What you will own
  • The HA management cluster (etcd and core services) on the management hardware.
  • Automated bare-metal provisioning for the GPU fleet: BMC/Redfish, PXE and virtual media, inspection, and lifecycle.
  • The NVIDIA GPU stack on Kubernetes through the GPU Operator (driver, container toolkit, device plugin, NFD/GFD, DCGM exporter), including whole-GPU allocation and per-tenant isolation.
  • The host side of the GPU fabric: NVLink and NVSwitch health through NVIDIA Fabric Manager, plus the RDMA NICs (ConnectX-8 class), jumbo frames, and GPUDirect setup on the node. The Network Engineer owns the switch fabric.
  • Secrets and certificate management (Vault or OpenBao, cert-manager). Identity, SIEM, and overall security posture will sit with a dedicated security hire.
  • GitOps-driven provisioning, so infrastructure changes are repeatable and auditable.
  • Secure tenant decommissioning: cryptographic NVMe wipe and GPU memory zeroing between tenants.
What We Are Looking For
Required
  • 5+ years in infrastructure or SRE, including 3+ years running Kubernetes in production with real node and cluster lifecycle ownership. On-prem or bare-metal experience counts for more here than managed-cloud Kubernetes.
  • Strong Linux fundamentals: systemd, kernel modules, NUMA, cgroups, NVMe and LVM storage, and host networking (bonding, VLANs, nftables). You are comfortable debugging at the hardware and driver level.
  • Hands-on server provisioning through IPMI, Redfish, or BMC tooling, with config management such as Ansible.
  • The NVIDIA GPU stack on Kubernetes: the GPU Operator and its components, GPU health and telemetry through DCGM, and GPU scheduling and per-tenant allocation.
  • Infrastructure as code and GitOps in production: Helm, Kustomize, and Argo CD or Flux.
  • Comfortable scripting and automating in Python and/or Go.
  • Self-directed: you find the problem, propose a fix, carry it out, and write it down.
Nice to have
  • Experience building to or operating NVIDIA reference architectures for AI compute (NCP, HGX).
  • A bare-metal provisioning system in production: Metal3/Ironic, MAAS, Tinkerbell, or Foreman.
  • Container security basics: RBAC, Pod Security Standards, network policies, and secrets management (Vault or OpenBao).
  • Distributed storage in production (Ceph/Rook or MinIO), or readiness to own it with vendor support.
  • KubeVirt with GPU passthrough for VM-level tenant isolation.
  • Awareness of the EU rules that shape this work: GDPR, the AI Act, DORA, NIS2, and NEN 7510
  • NVIDIA Run:ai or the KAI Scheduler for GPU scheduling and quota.
  • NVIDIA certifications (NCP-AIO, NCA-AIIO)
  • Experience at a GPU cloud, AI provider, or HPC centre.

We don't expect one person to be an expert in bare metal, GPUs, storage, and security all at once. The core we're hiring for is bare-metal Kubernetes, the NVIDIA GPU stack, and host-side GPU networking. The storage and security depth can be built up on the job, with support.

Why Join Us
  • Help build a sovereign EU GPU cloud from the ground up.
  • Own a critical platform layer, not just tickets or maintenance.
  • Work on modern AI infrastructure, GPU platforms, Kubernetes, observability, and automation.
  • Join a company with deep experience in OCP, datacenter, AI/HPC, and cloud infrastructure.
  • Build infrastructure for organizations where data location, compliance, and reliability truly matter.
  • Company gym.
  • Modern office in Hoofddorp.
  • Informal and open working culture.
  • Participation in relevant conferences and exhibitions across Europe.
  • Opportunity to develop your skills in a fast-growing technology environment.
Our Work Culture

Circle B offers an informal working atmosphere with energetic people who enjoy being part of a growing technology company. We have an open management culture and encourage colleagues to contribute to improving our products, services, and processes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer — AI Cloud
Senior Platform Engineer — AI Cloud

Circle B • Hoofddorp

On-site
EUR 70,000 - 90,000
Company gym
Participation in relevant conferences
Modern office in Hoofddorp
+1
Senior Network Engineer — AI / GPU Fabric
Senior Network Engineer — AI / GPU Fabric

Circle B • Hoofddorp

On-site
EUR 120,000 - 180,000
Office in Hoofddorp
Informal and open culture
Conference opportunities
Infrastructure Engineer
Infrastructure Engineer

Nebul • Leiden

On-site
EUR 55,000 - 75,000
Product Owner / Architect – Datacenter Infrastructure
Product Owner / Architect – Datacenter Infrastructure

Nebul Bv • Netherlands

Hybrid
EUR 110,000 - 140,000
Product Owner HPC Infrastructure
Product Owner HPC Infrastructure

Nebul Bv • Netherlands

Hybrid
EUR 90,000 - 130,000
Hybrid work
Equity options
Career growth opportunities
AI Cloud Platform Engineer: Observability & GPU Billing
AI Cloud Platform Engineer: Observability & GPU Billing

Circle B • Hoofddorp

On-site
EUR 70,000 - 90,000
Company gym
Participation in relevant conferences
Modern office in Hoofddorp
+1
Site Reliability Engineer – AI Cloud Platform
Site Reliability Engineer – AI Cloud Platform

Nebul • Leiden

On-site
EUR 90,000 - 120,000
Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Talanto • Amsterdam

Hybrid
EUR 90,000 - 130,000
Technical Project Manager (Hardware)
Technical Project Manager (Hardware)

Nebius • Amsterdam

On-site
EUR 90,000 - 130,000
Competitive pay
Career growth
Flexibility
+3
Senior Technical Program Manager - New Data Center Launches
Senior Technical Program Manager - New Data Center Launches

Nebius • Amsterdam

On-site
EUR 75,000 - 95,000
Competitive salary
Professional growth opportunities
Flexible working arrangements
+1