Senior Technical Support Engineer – Slurm

Jobtailor

Deutschland

Remote

EUR 70.000 - 100.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, versandbereit.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor is seeking a senior Slurm Administrator to own production AI and HPC clusters in a German-based environment. You will diagnose complex issues, optimize configurations, and guide customers to reliable recoveries.

Collaborating with engineering, you will create reproducible test cases, diagnostic tools, and training materials to strengthen Slurm expertise across the organization. This role emphasizes strong Linux administration and multitier problem solving.

Qualifikationen

  • 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments.
  • Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes
  • In-depth Linux system-administration and troubleshooting experience, including systemd, cgroups, authentication, networking, and database-backed services
  • Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources
  • Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation
  • Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity
  • Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform

Aufgaben

  • Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters
  • Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability
  • Solve Slurm configuration and policy issues involving partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints
  • Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required
  • Isolate problems across Slurm and dependencies including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems
  • Advise customers on Slurm configuration, upgrades, operational practices, system-resource management, and safe recovery from production incidents
  • Collaborate with engineering teams by producing technical descriptions, reproducible test cases, and defect reports
  • Develop guides, knowledge-base articles, diagnostic tools, and internal training to strengthen Slurm expertise across the support organization

Kenntnisse

Slurm Administration
Linux System Administration
Job Scheduling
Performance Diagnosis
Container Integration

Ausbildung

BS degree in Computer Science, Engineering, or related field

Tools

MUNGE
NVIDIA Base Command Manager
Bright Cluster Manager
Pyxis
Enroot
Apptainer
Singularity

Jobbeschreibung

  • Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters
  • Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability
  • Solve Slurm configuration and policy issues involving partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints
  • Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required
  • Isolate problems across Slurm and dependencies including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems
  • Advise customers on Slurm configuration, upgrades, operational practices, system-resource management, and safe recovery from production incidents
  • Collaborate with engineering teams by producing technical descriptions, reproducible test cases, and defect reports
  • Develop guides, knowledge-base articles, diagnostic tools, and internal training to strengthen Slurm expertise across the support organization
Requirements
  • BS degree in Computer Science, Engineering, or a related field, or equivalent experience
  • 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments, including business-critical outage incidents
  • Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes
  • Capacity to identify sophisticated Slurm incidents independently and guide them to a technically sound resolution
  • In-depth Linux system-administration and troubleshooting experience, including systemd, cgroups, authentication, networking, and database-backed services
  • Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources
  • Strong analytical and research skills, including distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems
  • Excellent written and verbal communication skills
  • Experience supporting large-scale Slurm environments containing thousands of nodes or GPUs
  • Experience diagnosing scheduler performance, job-throughput, controller-load, and database-scaling issues
  • Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation
  • Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity
  • Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform
Core Competencies

Demonstrates expert-level understanding of Slurm architecture and administration in production HPC and AI environments, with strong analytical skills for diagnosing complex issues and guiding resolutions. Proficient in Linux system administration and troubleshooting, with the ability to collaborate effectively and produce technical documentation.

Highest-signal resume keywords
  • Slurm Administration
  • Linux System Administration
  • Job Scheduling
  • Performance Diagnosis
  • Container Integration
Hard Skills
  • Slurm Configuration
  • Resource Management
  • Job Constraints
  • Systemd
  • Cgroups
  • Networking
  • Database Troubleshooting
  • HPC Integration
  • Performance Analysis
  • Source-Level Debugging
Soft Skills
  • Analytical Skills
  • Research Skills
  • Written Communication
  • Verbal Communication
Industry Keywords
  • AI Clusters
  • HPC Clusters
  • Production Environments
  • Multi-User Clusters
  • High Availability
Tools & Technologies
  • MUNGE
  • NVIDIA Base Command Manager
  • Bright Cluster Manager
  • Pyxis
  • Enroot
  • Apptainer
  • Singularity
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

HPC Infrastructure Engineer – GPU Clusters
HPC Infrastructure Engineer – GPU Clusters

Jobtailor • Deutschland

Hybrid
EUR 80.000 - 140.000
Software Golang Engineer (Slurm)
Software Golang Engineer (Slurm)

G-CORE INNOVATIONS SOCIETE A RESPONSABILITE LIMITEE • Deutschland

Hybrid
EUR 90.000 - 130.000
Competitive compensation
Flexible working hours
Hybrid or remote options
+5
Technical Lead – GPU Infrastructure
Technical Lead – GPU Infrastructure

Jobtailor • Deutschland

Remote
EUR 120.000 - 160.000
Software Platform Support Engineer – GPU Cloud
Software Platform Support Engineer – GPU Cloud

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
HPC Systems Administrator
HPC Systems Administrator

EngineersOfAI • München

Vor Ort
EUR 60.000 - 80.000
Compute Solution Architect
Compute Solution Architect

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Senior GPU Cloud, K8S Expert
Senior GPU Cloud, K8S Expert

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
SRE Monitoring Platform Software Engineer, Entry Level
SRE Monitoring Platform Software Engineer, Entry Level

Jobtailor • Deutschland

Remote
EUR 42.000 - 64.000
Senior Technical Support Engineer
Senior Technical Support Engineer

Jobtailor • Deutschland

Remote
EUR 103.000 - 129.000
Staff Engineer
Staff Engineer

Jobtailor • Deutschland

Vor Ort
EUR 120.000 - 180.000