Infrastructure Engineer

Nebul

Leiden

On-site

EUR 55,000 - 75,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading cloud solutions provider is seeking an Infrastructure Engineer specializing in GPU Datacenter & Kubernetes. This full-time role in Leiden involves deploying and maintaining GPU-optimized Kubernetes clusters, integrating NVIDIA GPU Operators, and optimizing resource utilization. The ideal candidate has significant experience with Kubernetes in production, strong knowledge of NVIDIA technologies, and scripting skills in Python or Go. Applicants must reside in the Netherlands and have a valid work permit.

Qualifications

  • Experience running Kubernetes clusters in production, especially for GPU workloads.
  • Familiarity with NVIDIA GPU integration and scheduling logic.
  • Good programming/scripting skills in Go or Python.

Responsibilities

  • Deploy and maintain Kubernetes clusters optimized for GPU workloads.
  • Integrate and manage NVIDIA GPU Operators and plugins.
  • Support performance tuning, incident response, and upgrades.

Skills

Kubernetes
NVIDIA GPU integration
Linux
networking
Python
Go

Tools

Terraform
Ansible
Helm
Prometheus
Grafana

Job description

Direct message the job poster from Nebul

People Power Business! Mastering Talent, Recruitment, and Culture Integration.

Join Nebul

Nebul is a leader in sovereign‑hybrid cloud solutions, combining the security of private cloud infrastructure with the scalability of global hyperscalers. Rooted in European values of privacy, security, and compliance, Nebul enables businesses to harness AI with confidence. Our infrastructure powers large‑scale AI workloads, simulations, and real‑time analytics. If working on high‑performance infrastructure excites you, this is your opportunity.

What You’ll Be Doing

As an Infra Engineer – GPU Datacenter & Kubernetes, you’ll be central to building and operating Nebul’s performance‑critical infrastructure for AI. You will:

  • Deploy, maintain, and scale Kubernetes clusters optimized for GPU workloads
  • Integrate and manage NVIDIA GPU Operators, plugins, MIG configurations, and GPU scheduling logic
  • Automate infrastructure (compute, storage, networking) with Terraform, Ansible, Helm or equivalent tools
  • Optimize GPU resource utilization, minimize fragmentation, and ensure high throughput
  • Instrument clusters with observability tooling (Prometheus, DCGM, Grafana, OpenTelemetry)
  • Ensure secure multi‑tenant usage, RBAC, network policies, and isolation between workloads
  • Collaborate with AI/ML, product, and security teams to understand workload needs and align infrastructure strategy
  • Evolve architecture over time and contribute to platform roadmap and design decisions
  • Support performance tuning, incident response, capacity planning, and upgrades
  • Optionally, mentor others or lead small infrastructure projects or pods as the platform matures
Key Responsibilities
  • Architect and operate GPU‑accelerated Kubernetes clusters with high availability and performance
  • Build and maintain custom controllers, operators, or scheduling extensions to support NVIDIA features
  • Enforce multi‑tenant security and resource isolation via RBAC, namespaces, network policies, and policy engines
  • Monitor GPU & cluster health; build dashboards, alerts, telemetry, and feedback loops
  • Tune systems, detect bottlenecks, and iterate on GPU, networking, storage performance
  • Drive capacity planning and cost efficiency
  • Evaluate and integrate new GPU and infrastructure technologies
  • Document architecture, operational playbooks, and best practices
What You Bring
  • Solid experience in running Kubernetes clusters in production, especially for GPU workloads
  • Deep familiarity with NVIDIA GPU integration: GPU Operator, device plugins, MIG, scheduling logic
  • Proficiency with Linux, networking, container runtimes, and GPU toolkits
  • Experience in monitoring, telemetry, and observability stacks
  • Understanding of AI/ML or HPC workload patterns and how they stress infrastructure
  • Good programming/scripting skills in Go or Python (for tooling, controllers)
  • Ownership mindset, reliability focus, strong communication, ability to work across functions
Bonus Points If You Have
  • Kernel or driver‑level GPU/compute knowledge
  • Experience with schedulers like Slurm, Volcano, or custom scheduling plugins
  • Contributions to open‑source infrastructure or GPU‑Kubernetes ecosystem
  • Experience with hybrid‑cloud or multi‑cloud GPU orchestration
  • Exposure to storage optimization for AI workloads (NVMe, Lustre, Ceph, NVMf)
Eligibility & Application Information
  • We welcome non‑native Dutch speakers. To apply:
  • Valid work permit in the Netherlands
  • Reside in the Netherlands and be able to travel to the office near The Hague

Nebul does not offer relocation assistance or sponsorship.

Referrals increase your chances of interviewing at Nebul by 2x.

Mid‑Senior level | Full‑time | Data Infrastructure and Analytics, IT System Custom Software Development

Leiden, South Holland, Netherlands

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer – AI Cloud Platform
Site Reliability Engineer – AI Cloud Platform

Nebul • Leiden

On-site
EUR 70,000 - 110,000
Cloud Engineer
Cloud Engineer

Nebius • Amsterdam

On-site
EUR 60,000 - 80,000
Technical Project Manager (Hardware)
Technical Project Manager (Hardware)

Nebius • Amsterdam

On-site
EUR 90,000 - 130,000
Competitive pay
Career growth
Flexibility
+3
DevOps Engineer – Go & Cloud Platform
DevOps Engineer – Go & Cloud Platform

Nebul • Leiden

On-site
EUR 75,000 - 110,000
Data Center - Service Delivery Manager
Data Center - Service Delivery Manager

Nebius • Amsterdam

On-site
EUR 75,000 - 95,000
Competitive salary
Comprehensive benefits package
Opportunities for professional growth
+1
Platform Engineer
Platform Engineer

Nebul • Leiden

On-site
EUR 60,000 - 100,000
Engineering Team Lead
Engineering Team Lead

Nebul • Leiden

On-site
EUR 90,000 - 120,000
GPU-Driven Kubernetes Infra Engineer for AI Workloads
GPU-Driven Kubernetes Infra Engineer for AI Workloads

Nebul • Leiden

On-site
EUR 55,000 - 75,000
Senior Technical Program Manager - New Data Center Launches
Senior Technical Program Manager - New Data Center Launches

Nebius • Amsterdam

On-site
EUR 75,000 - 95,000
Competitive salary
Professional growth opportunities
Flexible working arrangements
+1
Technical Product Manager – AI Compute Platform
Technical Product Manager – AI Compute Platform

ApplyMint • Netherlands

Hybrid
EUR 70,000 - 100,000
Competitive compensation
Career growth opportunities
Flexible work environment