Direct message the job poster from Nebul
People Power Business! Mastering Talent, Recruitment, and Culture Integration.
Join Nebul
Nebul is a leader in sovereign‑hybrid cloud solutions, combining the security of private cloud infrastructure with the scalability of global hyperscalers. Rooted in European values of privacy, security, and compliance, Nebul enables businesses to harness AI with confidence. Our infrastructure powers large‑scale AI workloads, simulations, and real‑time analytics. If working on high‑performance infrastructure excites you, this is your opportunity.
What You’ll Be Doing
As an Infra Engineer – GPU Datacenter & Kubernetes, you’ll be central to building and operating Nebul’s performance‑critical infrastructure for AI. You will:
- Deploy, maintain, and scale Kubernetes clusters optimized for GPU workloads
- Integrate and manage NVIDIA GPU Operators, plugins, MIG configurations, and GPU scheduling logic
- Automate infrastructure (compute, storage, networking) with Terraform, Ansible, Helm or equivalent tools
- Optimize GPU resource utilization, minimize fragmentation, and ensure high throughput
- Instrument clusters with observability tooling (Prometheus, DCGM, Grafana, OpenTelemetry)
- Ensure secure multi‑tenant usage, RBAC, network policies, and isolation between workloads
- Collaborate with AI/ML, product, and security teams to understand workload needs and align infrastructure strategy
- Evolve architecture over time and contribute to platform roadmap and design decisions
- Support performance tuning, incident response, capacity planning, and upgrades
- Optionally, mentor others or lead small infrastructure projects or pods as the platform matures
Key Responsibilities
- Architect and operate GPU‑accelerated Kubernetes clusters with high availability and performance
- Build and maintain custom controllers, operators, or scheduling extensions to support NVIDIA features
- Enforce multi‑tenant security and resource isolation via RBAC, namespaces, network policies, and policy engines
- Monitor GPU & cluster health; build dashboards, alerts, telemetry, and feedback loops
- Tune systems, detect bottlenecks, and iterate on GPU, networking, storage performance
- Drive capacity planning and cost efficiency
- Evaluate and integrate new GPU and infrastructure technologies
- Document architecture, operational playbooks, and best practices
What You Bring
- Solid experience in running Kubernetes clusters in production, especially for GPU workloads
- Deep familiarity with NVIDIA GPU integration: GPU Operator, device plugins, MIG, scheduling logic
- Proficiency with Linux, networking, container runtimes, and GPU toolkits
- Experience in monitoring, telemetry, and observability stacks
- Understanding of AI/ML or HPC workload patterns and how they stress infrastructure
- Good programming/scripting skills in Go or Python (for tooling, controllers)
- Ownership mindset, reliability focus, strong communication, ability to work across functions
Bonus Points If You Have
- Kernel or driver‑level GPU/compute knowledge
- Experience with schedulers like Slurm, Volcano, or custom scheduling plugins
- Contributions to open‑source infrastructure or GPU‑Kubernetes ecosystem
- Experience with hybrid‑cloud or multi‑cloud GPU orchestration
- Exposure to storage optimization for AI workloads (NVMe, Lustre, Ceph, NVMf)
Eligibility & Application Information
- We welcome non‑native Dutch speakers. To apply:
- Valid work permit in the Netherlands
- Reside in the Netherlands and be able to travel to the office near The Hague
Nebul does not offer relocation assistance or sponsorship.
Referrals increase your chances of interviewing at Nebul by 2x.
Mid‑Senior level | Full‑time | Data Infrastructure and Analytics, IT System Custom Software Development
Leiden, South Holland, Netherlands