HPC Cluster Architect

nexgencloud

Deutschland

Hybrid

EUR 120.000 - 180.000

Vollzeit

Vor 12 Tagen
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Annual discretionary bonus
25 days holiday
Remote or hybrid options
Wellbeing benefits

Zusammenfassung

nexgencloud is seeking a Senior HPC Cluster Architect to own the design and delivery of large‑scale GPU compute clusters. You will guide end‑to‑end architecture across compute, networking, storage and physical design, turning customer requirements into production deployments.

You’ll act as a hands‑on technical authority with deep HPC experience, driving standardised reference designs and collaborating with Solutions Architecture, Cloud Engineering and data centre partners to bring designs to

Qualifikationen

  • Proven experience designing and delivering HPC or AI software stacks at scale.

Aufgaben

  • Own end‑to‑end cluster architecture for large GPU deployments from requirements to rack layouts and production handover.

Tools

Docker
Kubernetes
NVIDIA GPU Operator
NCCL tools

Jobbeschreibung

Role overview

A senior HPC Cluster Architect role focused on designing and delivering large‑scale GPU compute clusters. The position owns end‑to‑end architecture across compute, networking, storage, and physical design, translating customer requirements into production‑ready, commercially optimised deployments. It is a hands‑on technical authority role for someone with deep HPC experience who wants to see designs go live.

Responsibilities
  • Own end‑to‑end cluster architecture for large‑scale NVIDIA GPU deployments, from initial customer requirements through rack layouts, BOM, power and cooling design, and production handover
  • Design high‑performance network fabrics spanning compute (InfiniBand, RDMA, NVLink/NVSwitch), storage, and WAN, including topology, oversubscription models, and scaling strategies
  • Engage directly with OEMs and vendors to validate hardware configurations, review quotes, and balance technical soundness with commercial optimisation
  • Provide technical oversight during deployment and bring‑up, supporting hardware validation, performance testing, and acting as escalation point for complex integration issues
  • Act as a senior technical leader across Solutions Architecture, Cloud Engineering, and data centre partners, contributing to standardised reference designs and growing the HPC engineering function
Requirements
  • Proven experience designing and delivering HPC or AI software stacks at scale, including workload profiling, scheduler configuration (SLURM, PBS, or equivalent), MPI/NCCL tuning, and distributed training frameworks such as PyTorch, JAX, or DeepSpeed
  • Deep understanding of GPU software environments, including CUDA, cuDNN, NCCL, driver stacks, and the tooling needed to run large‑scale AI training and inference reliably in production
  • Hands‑on experience optimising AI and HPC workloads across multi‑GPU and multi‑node setups, covering profiling, bottleneck identification, and performance tuning at both application and infrastructure layers
  • Working knowledge of containerisation and orchestration in HPC/AI contexts: Docker, Kubernetes, NVIDIA GPU Operator, and container‑native workload management
  • Background in an OEM, hyperscaler, neo-cloud, or enterprise/research HPC environment, with exposure to the full design‑to‑deployment lifecycle for GPU‑accelerated workloads
  • Ability to produce clear technical documentation and architecture diagrams for both engineering and executive audiences, with confidence engaging customers, vendors, and internal teams as a technical authority
Nice to have
  • Experience with large‑scale cluster performance benchmarking (NCCL tests, MLPerf, or equivalent) and familiarity with expected outcomes across GPU generations and topologies
  • Exposure to MLOps tooling and AI platform layers, including experiment tracking (MLflow, W&B), model serving frameworks (Triton, vLLM), and pipeline orchestration (Kubeflow, Airflow)
  • Familiarity with InfiniBand and high‑performance networking as it relates to distributed training performance, sufficient to engage credibly on topology and tuning decisions
Benefits and work setup
  • Competitive salary with an annual discretionary bonus scheme
  • Employee wellbeing benefits and 25 days of holiday plus public holidays
  • Flexible working arrangements, with remote or hybrid options depending on role and location
  • Real ownership and autonomy, with the trust to take initiative and experiment
  • Clear career progression and growth opportunities within a fast‑growing organisation
  • Collaborative, international culture built on trust, transparency, and ownership
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior HPC GPU Cluster Lead for Deep Learning Infra
Senior HPC GPU Cluster Lead for Deep Learning Infra

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063
Infrastructure Engineer (GPU & Compute)
Infrastructure Engineer (GPU & Compute)

lightningai • Deutschland

Hybrid
EUR 155.000 - 189.000
Discretionary bonus
Equity
Comprehensive medical coverage
+3
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 140.000
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063
Technical Lead – GPU Infrastructure
Technical Lead – GPU Infrastructure

Jobtailor • Deutschland

Remote
EUR 120.000 - 160.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 170.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Corporation • Berlin

Vor Ort
EUR 120.000 - 180.000
HPC Infrastructure Engineer – GPU Clusters
HPC Infrastructure Engineer – GPU Clusters

Jobtailor • Deutschland

Hybrid
EUR 80.000 - 140.000
Senior HPC Engineer, GPU Compute
Senior HPC Engineer, GPU Compute

Meyandy LLC • Berlin

Hybrid
EUR 120.000 - 180.000
Head of Compute Engineering
Head of Compute Engineering

Impossible Cloud GmbH • Hamburg

Vor Ort
EUR 80.000 - 100.000
Competitive salary
ESOP
Subsidized gym membership
+1