HPC Systems Engineer – Security Architecture & Cluster Administration

Cyber Valley GmbH

Tübingen

On-site

EUR 65,000 - 90,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Flexible working hours
Hybrid/remote option
Professional development & conferences

Job summary

Cyber Valley GmbH in Germany seeks an HPC Systems Engineer to shape security architecture and manage production ML HPC environments across multiple data centers. You will harden clusters, design secure provisioning, and lead incident response for researchers and compute workloads.

The role blends hands-on cluster operations with security leadership, offering flexible hours and an international team, with opportunities to influence the roadmap and professional growth.

Qualifications

  • Masters degree in Computer Science or a related field.
  • Hands-on HPC background with Slurm, parallel file systems, and high-speed networks.
  • Experience with virtualization for management and infrastructure services.
  • Strong scripting and automation skills; IaC experience.

Responsibilities

  • Design and operate HPC clusters across four data centers (SLURM, filesystems, networks).
  • Conceive and establish security architecture for the ML Cloud and harden the HPC environment.
  • Automate provisioning and config of heterogeneous compute/storage/network nodes.
  • Run patch and vulnerability management across heterogeneous systems.
  • Build and operate logging, monitoring and IDS; integrate telemetry into dashboards and incident-response workflows.
  • Lead incident response for clusters: detection, containment, forensic support, post-incident review.
  • Automate operations and security policy as code (Ansible/IaC).
  • Advise researchers on secure cluster usage and derive roadmap requirements.

Skills

Threat modelling
Scripting (Bash, Python)
Incident response
English communication
Collaborative mindset

Education

Master's degree in Computer Science or related field

Tools

Proxmox
SLURM
Weka
Lustre
Ceph
Ansible
Docker

Job description

HPC Systems Engineer – Security Architecture & Cluster Administration

The Cluster of Excellence "Machine Learning - New Perspectives for Science" together with the Tübingen AI Center at the University of Tübingen offers a position as

HPC Systems Engineer – Security Architecture & Cluster Administration (m/f/d, E13 TV-L, 100%)

The position is available in the team of the Machine Learning Science Cloud and runs until 31st December 2032.

About Them

The ML Cloud team designs, runs and safeguards high-performance computing infrastructure tailored to machine learning workloads. Their state of the art systems are distributed across four data centers for the Tübingen AI Center, the Cluster of Excellence ’Machine Learning: New Perspectives for Science’ and the Hertie Institute for AI in Brain Health. Every day, researchers unleash thousands of compute jobs on their systems, training frontier-scale neural networks and running experiments.

The Opportunity

Traditionally, HPC environments prioritized only performance and open access, with security as an afterthought. As their clusters grow and handle increasingly sensitive research data, security becomes mission-critical. They are looking for an HPC System Engineer who combines hands-on cluster operations with a security mindset. You will actively shape the security architecture of a live, production HPC / ML environment, keep their systems hardened, and build the processes and infrastructure that protect researchers, their data and their compute against a fast-evolving threat landscape.

Your Tasks
  • Design and operate theirHPC clusters across four data centers, including scheduler (SLURM), parallel filesystems, networks and accelerators, ensuring high availability and throughput for research workloads
  • Conceive and establish the security architecture of the Machine Learning Science Cloud and harden the HPC environment
  • Evolve the automated provisioning and configuration of heterogeneous compute, storage and network nodes (e.g. image-based provisioning, node lifecycle)
  • Run patch and vulnerability management - risk assessment across heterogeneous systems
  • Build and operate logging, monitoring and intrusion detection, and integrate HPC telemetry into both operational dashboards and incident-response workflows
  • Lead incident response for the clusters: detection, containment, forensic support and post-incident review
  • Automate operations and security policy as code (Ansible/IaC)
  • Advice researchers on efficient, secure cluster usage (job scheduling, data handling, access workflows) and derive requirements for our further roadmap from their scientific workloads
Your Profile
  • Masters degree in Computer Science or a related field
  • In-depth IT security knowledge: system hardening, network security, applied cryptography and IAM - with the ability to derive architectural decisions from a threat model, not only to apply given baselines
  • Hands‑on HPC background: Slurm, parallel file systems (Weka, Lustre, Ceph), GPU workloads and high‑speed networks (InfiniBand, 400G Ethernet)
  • Experience with virtualization for management‑plane and infrastructure services (Proxmox)
  • Strong scripting and automation skills (Bash, Python, Ansible) and experience with configuration management / Infrastructure‑as‑Code
  • A plus: security frameworks (ISO 27001, BSI Grundschutz), container security (Apptainer/Singularity/Docker) or offensive‑security fundamentals
  • Independent, structured working style and good communication in English; German is a plus
  • A collaborative, user‑facing mindset – comfortable supporting and advising researchers and translating their needs into platform design
What They Offer
Technically deep, architecturally open work that directly enables cutting‑edge machine‑learning research
  • Flexible working hours and the option to work partially from home
  • Working in an English‑speaking, international team of HPC experts
  • A small, senior team with flat hierarchy where responsibility is split by domain
  • Ownership of a technical domain in a production environment of real scale
  • Professional development, conference attendance and real influence on our roadmap
Interested? They'd love to hear from you!

For questions, please contact Simon Kreuzer atsimon.kreuzer@uni-tuebingen.de . The position is available from now on. Hiring is done by the Central Administration of the University of Tübingen. Severely disabled persons will be given preferential consideration if equally qualified. The University of Tübingen is committed to equity and diversity and actively promotes equal opportunities. The position is divisible.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Systems Administrator
HPC Systems Administrator

EngineersOfAI • München

On-site
EUR 60,000 - 80,000
Senior Platform Engineer AI & Observability (m/f/d)
Senior Platform Engineer AI & Observability (m/f/d)

LRZ • Garching bei München

On-site
EUR 70,000 - 90,000
Up to 60% mobile work
Free parking
Pension plan (VBL)
+2
PhD position in Scientific Machine Learning: Data science at scale and mixed precision solvers
PhD position in Scientific Machine Learning: Data science at scale and mixed precision solvers

Technische Universität München (Technical University of Munich) • München

On-site
EUR 50,000 - 70,000
HPC Systems Administrator
HPC Systems Administrator

Helsing • München

On-site
EUR 55,000 - 75,000
Competitive salary
Relocation support
Health & wellness support
+2
Senior Specialist Scientific IT (f/m/x)
Senior Specialist Scientific IT (f/m/x)

DZNE - German Center for Neurodegenerative Diseases • Bonn

Hybrid
EUR 60,000 - 90,000
Family service
Health promotion programs
System Administrator / System Technician for Cybersecurity
System Administrator / System Technician for Cybersecurity

Fraunhofer Gesellschaft • Darmstadt

On-site
EUR 60,000 - 85,000
Flexible Arbeitszeitmodelle
Familienfreundliches Arbeitsumfeld
Betriebliche Altersvorsorge
+1
HPC Linux System Administrator
HPC Linux System Administrator

Global Market Solutions • München

Hybrid
EUR 70,000 - 90,000
Senior AI Platform & Research Infrastructure Engineer (m/f/d)
Senior AI Platform & Research Infrastructure Engineer (m/f/d)

LRZ • Garching bei München

On-site
EUR 110,000 - 170,000
60% remote work option
Free parking
Pension plan (VBL)
+1
HPC-, Data- und KI-Systemmanager:in (w/m/d)
HPC-, Data- und KI-Systemmanager:in (w/m/d)

GoHiring GmbH • Darmstadt

On-site
EUR 60,000 - 80,000
30 Tage Urlaub
Mobiles Arbeiten
VBL Zusatzversorgung
+2
PhD Position in Trustworthy Biomedical AI (f/m/x)
PhD Position in Trustworthy Biomedical AI (f/m/x)

LungResearch@HelmholtzMunich • München

On-site
EUR 30,000 - 42,000
30 days annual leave
Flexi days
Public holidays
+1