Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure

Leibniz-Rechenzentrum der Bayerischen Akademie der Wissenschaften

Garching bei München

Hybrid

EUR 60.000 - 85.000

Vollzeit

Vor 5 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Mobile work up to 60%
Free parking
Pension plan (VBL)
Modern equipment
Discounted gym membership

Zusammenfassung

Leibniz-Rechenzentrum der Bayerischen Akademie der Wissenschaften in Bayern is seeking a Systems Engineer/Administrator to plan, deploy, and operate an expanding AI service infrastructure for researchers within a dedicated team.

You will run day-to-day operations of high-performance AI/Data Analytics clusters, apply best practices in automation and monitoring, and collaborate with providers and partners to maintain secure, scalable systems.

Qualifikationen

  • Master’s degree or equivalent with research/enterprise admin experience.
  • Experience operating high‑performance clusters for AI/Data Analytics.
  • Proficiency in Linux, Git, Gitlab and data center environments.
  • Knowledge of job scheduling, container orchestration, networking and storage.
  • Experience with automation tools and scripting (Bash, Python, Perl).
  • Strong English writing/presentation; German preferred.

Aufgaben

  • Plan, deploy, and operate AI service infrastructure for researchers.
  • Run day‑to‑day operations of AI/Analytics HPC clusters.
  • Collaborate with hardware/software providers and partners.
  • Develop and implement certified processes and workflows.

Kenntnisse

Cluster administration
AI infrastructure
Linux system administration
Automation & monitoring
Collaboration & communication

Ausbildung

Master’s degree in CS or scientific computing
Bachelor’s degree with >3 years admin/engineering experience

Tools

Slurm Workload Manager
Kubernetes
GPFS/NFS
InfiniBand networking
LDAP/Shibboleth
Ansible / IaC

Jobbeschreibung

Description:

Looking for an employer you can count on? Join us!

We are looking to expand our team and are hiring:

Systems Engineer/Administrator(m/f/d) for AI Service Infrastructure

Your Role and Responsibilities:
  • Planning, deploying and operating an ever-growing AI service infrastructure for researchers and academic users as a member of a dedicated team within a friendly and open work environment
  • Running the day-to-day operations of high-performance cluster systems focused on AI and Data Analytics applications & workflows, following best-practices of system administration, automation and monitoring
  • Interacting with technology providers and developers (hardware/software) as well as national and European collaboration partners and projects
  • You actively contribute to the development and implementation of certified and transparent processes.
Your Qualifications:

Required/Minimum Qualifications

  • Master’s degree in computer science or other areas of scientific computing (e.g. physics, math), or Bachelor’s degree with multiple years (>3) of practical experience in system administration/engineering in research or enterprise environments

Other Requirements

  • Proficiency working in data center environments (incl. Linux, Git, Gitlab)
  • Advanced knowledge and experience in the administration/management of high-performance cluster systems for AI and Data Analytics, covering the majority of the following areas:
    • Compute resources, including
      • Accelerated server compute nodes (incl. GPUs)
      • Job scheduling and resource management (with Slurm Workload Manager)
      • Container orchestration (with Kubernetes)
    • Low latency fabrics (incl. InfiniBand), Ethernet and additional networking technologies (DNS, Firewall, etc.)
    • Fabric attached high-performance storage (e.g. GPFS, NFS, RDMA)
    • Authentication/Authorization (incl. LDAP, Shibboleth)
  • Practical knowledge in architecting/designing, deploying and maintaining hardware and software systems based on components/technologies as listed above
  • Experience using automated tools for deployment and configuration management (i.e. infrastructure as code, e.g. Ansible), but not afraid to script on CLI, either (e.g. Bash, Python, Perl)
  • Knowledge and experience in software deployment and software lifecycle management (ideally based on principles of continuous integration/continuous deployment, CI/CD)
  • Willingness to keep up with current developments and to learn new technologies in the field of AI
  • Friendly contact with colleagues/collaboration partners, and, where necessary, with users in higher level support
  • Excellent English and preferably German writing and presentation skills
  • You like to share your experience and review existing processes to identify opportunities for improvement within the team.
Of course, we provide comprehensive onboarding for career starters and professionals returning to the workforce.
Additional or Preferred Qualifications:
  • Knowledge of machine learning/AI applications, frameworks and methods
  • Experience with IT security management within an ISO 27001-compliant framework, including process definition, risk management, release and deployment management, and patching
Conditions:

Working time

full time (40 hrs) / part-time possible

flexible working model with electronic time recording

Contract duration

24 months, a further employment is intended

Remuneration

up to E 13, seeEntgelttabelle TV-L

Annual leave / compensatory time off

30 days

(24.12. + 31.12. additionally day off)

Overtime is compensated by additional time off

Benefits

Up to 60% mobile work for suitable tasks, mostly two-person offices, bus and subway (U6) on the doorstep, free parking, pension plan of the Versorgungsanstalt des Bundes und der Länder (VBL), state-of-the-art work equipment, Cooperation Agreement with Sports Facility Providers for Discounted Membership

What you can expect at LRZ:

Are you looking for a multifaceted and intellectually stimulating position in a dynamic, cooperative and innovative work environment? Then LRZ is the place to be for you! Here at LRZ a collegial, appreciative work environment meets an international crowd of experts who work together to advance IT services for ground-breaking research. We offer flexible work schemes for an optimal work-life balance. Our staff values their creative leeway. As an institute of the Bavarian Academy of Science and Humanities we offer all the benefits of public service. And of course, no wishes remain unfulfilled at the LRZ in terms of technical equipment. We share experiences, constantly review and improve our processes, and are proud that our service-quality and data-security are regularly certified and rated highly.

We actively promote diversity and welcome applications from talented individuals, regardless of cultural background, nationality, ethnicity, gender and sexual identity, physical abilities, religion and age. We give priority to applications from people with disabilities who are equally qualified (SGB IX).

Are you unsure whether the job suits you or you suit us? Or do you still have questions about this position? Our colleagues will be happy to answer all your questions.

Hereyou will find information about the collection of personal data during the application process.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure
Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure

Leibniz-Rechenzentrum der Bayer. Akad. d. Wissenschaften • Garching bei München

Hybrid
EUR 75.000 - 105.000
Up to 60% remote work
Free parking
Pension plan (VBL)
+2
System Engineer (m/w/d) Open Source Collaboration & Kubernetes
System Engineer (m/w/d) Open Source Collaboration & Kubernetes

Leibniz-Rechenzentrum der Bayerischen Akademie der Wissenschaften • Garching bei München

Vor Ort
EUR 65.000 - 90.000
JobRad
Parkplatz
Weiterbildung
System Engineer (m/w/d) Open Source Collaboration & Kubernetes
System Engineer (m/w/d) Open Source Collaboration & Kubernetes

Leibniz-Rechenzentrum • Garching bei München

Hybrid
EUR 60.000 - 75.000
JobRad
Parkplatz
VBL Altersvorsorge
+1
Senior AI Expert (m/f/d) for AI Computing Services
Senior AI Expert (m/f/d) for AI Computing Services

Leibniz-Rechenzentrum • Garching bei München

Hybrid
EUR 65.000
Pension plan
State-of-the-art work equipment
Free parking
System Engineer (m/w/d) Open Source Collaboration & Kubernetes
System Engineer (m/w/d) Open Source Collaboration & Kubernetes

Meyandy LLC • Garching bei München

Hybrid
EUR 70.000 - 95.000
40 Std/Woche
60% remote möglich
TV-L Entgelt
Assistenz (m/w/d) für Führungskräfte im Bereich Forschung | Teilzeit 20 Std
Assistenz (m/w/d) für Führungskräfte im Bereich Forschung | Teilzeit 20 Std

Leibniz-Rechenzentrum • Garching bei München

Hybrid
EUR 27.000 - 38.000
Homeoffice bis zu 60%
Jobticket
Parkplatz
+3
Wissenschaftliche Beschäftigte (m/w/d) zur Administration und Weiterentwicklung von Enterprise [...]
Wissenschaftliche Beschäftigte (m/w/d) zur Administration und Weiterentwicklung von Enterprise [...]

Leibniz-Rechenzentrum • Garching bei München

Hybrid
EUR 52.000 - 78.000
30 Tage Urlaub
Individuelle Weiterbildungen
Mobiles Arbeiten bis zu 60%
+2
Personalsachbearbeiterin (m/w/d)
Personalsachbearbeiterin (m/w/d)

Leibniz-Rechenzentrum der Bayerischen Akademie der Wissenschaften • Garching bei München

Vor Ort
EUR 60.000 - 75.000
30 Tage Urlaub
Flexible Arbeitszeiten
Mobiles Arbeiten
+2
Netzwerk-Architekt (w/m/d)
Netzwerk-Architekt (w/m/d)

Deutsches Zentrum für Luft- und Raumfahrt e. V. • Berlin

Vor Ort
EUR 70.000 - 100.000
Fort- und Weiterbildungsmöglichkeiten
Vielfalt & Gleichstellung
Bewerbungen schwerbehinderter Menschen
Informatiker:in (w/m/d)
Informatiker:in (w/m/d)

Forschungszentrum Jülich • Jülich

Hybrid
EUR 50.000 - 70.000
30 Tage Urlaub
Betriebliches Gesundheitsmanagement
Flexible Arbeitszeitmodelle
+1