Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure

Leibniz-Rechenzentrum der Bayer. Akad. d. Wissenschaften

Garching bei München

Hybrid

EUR 75.000 - 105.000

Vollzeit

Vor 13 Tagen
Bewerbungsgenerator

Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Up to 60% remote work
Free parking
Pension plan (VBL)
State-of-the-art equipment
Sports facility discounts

Zusammenfassung

Leibniz-Rechenzentrum der Bayer. Akad. d. Wissenschaften sucht einen Systems Engineer/Administrator (m/f/d) für AI-Service-Infrastruktur.

Sie planen, implementieren und betreiben eine wachsende AI-Service-Infrastruktur für Forscher und akademische Nutzer in einem freundlichen Umfeld. Sie arbeiten an HPC-Clustern, führen automatisierung und Monitoring durch und arbeiten eng mit Partnern zusammen. Sie verfügen über einen Master in Informatik oder vergleichbare Qualifikation und besitzen fundierte

Qualifikationen

  • Master’s degree in computer science or related field, or equivalent experience.
  • Extensive experience in system administration/engineering in research or enterprise environments.
  • Proficiency with Linux, Git, Gitlab, and HPC cluster management.

Aufgaben

  • Plan, deploy, and operate AI service infrastructure for researchers.
  • Run day-to-day operations of high-performance cluster systems for AI and data analytics.
  • Collaborate with hardware/software providers and national/European partners.
  • Contribute to certified, transparent processes and best practices.

Kenntnisse

Linux
Git
Gitlab
Kubernetes
Slurm
Ansible
Bash
Python
CI/CD
Networking
InfiniBand

Ausbildung

Master’s degree in computer science or related field
Bachelor’s degree with 3+ years of experience

Tools

Kubernetes
Slurm Workload Manager
GPFS/NFS
LDAP/Shibboleth
Ansible
Gitlab
InfiniBand Networking

Jobbeschreibung

Looking for an employer you can count on? Join us!

Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure
Your Role and Responsibilities
  • Planning, deploying and operating an ever-growing AI service infrastructure for researchers and academic users as a member of a dedicated team within a friendly and open work environment
  • Running the day-to-day operations of high-performance cluster systems focused on AI and Data Analytics applications & workflows, following best-practices of system administration, automation and monitoring
  • Interacting with technology providers and developers (hardware/software) as well as national and European collaboration partners and projects
  • You actively contribute to the development and implementation of certified and transparent processes.
Required/Minimum Qualifications
  • Master’s degree in computer science or other areas of scientific computing (e.g. physics, math), or Bachelor’s degree with multiple years (>3) of practical experience in system administration/engineering in research or enterprise environments
Other Requirements
  • Proficiency working in data center environments (incl. Linux, Git, Gitlab)
  • Advanced knowledge and experience in the administration/management of high-performance cluster systems for AI and Data Analytics, covering the majority of the following areas:
  • Compute resources, including
  • Accelerated server compute nodes (incl. GPUs)
  • Job scheduling and resource management (with Slurm Workload Manager)
  • Container orchestration (with Kubernetes)
  • Low latency fabrics (incl. InfiniBand), Ethernet and additional networking technologies (DNS, Firewall, etc.)
  • Fabric attached high-performance storage (e.g. GPFS, NFS, RDMA)
  • Authentication/Authorization (incl. LDAP, Shibboleth)
  • Practical knowledge in architecting/designing, deploying and maintaining hardware and software systems based on components/technologies as listed above
  • Experience using automated tools for deployment and configuration management (i.e. infrastructure as code, e.g. Ansible), but not afraid to script on CLI, either (e.g. Bash, Python, Perl)
  • Knowledge and experience in software deployment and software lifecycle management (ideally based on principles of continuous integration/continuous deployment, CI/CD)
  • Willingness to keep up with current developments and to learn new technologies in the field of AI
  • Friendly contact with colleagues/collaboration partners, and, where necessary, with users in higher level support
  • Excellent English and preferably German writing and presentation skills
  • You like to share your experience and review existing processes to identify opportunities for improvement within the team.
Additional or Preferred Qualifications
  • Knowledge of machine learning/AI applications, frameworks and methods
  • Experience with IT security management within an ISO 27001-compliant framework, including process definition, risk management, release and deployment management, and patching
Conditions
  • Working time: full time (40 hrs) / part-time possible, flexible working model with electronic time recording
  • Contract duration: 24 months, a further employment is intended
  • Remuneration: up to E 13, see Entgelttabelle TV-L
  • Annual leave / compensatory time off: 30 days (24.12. + 31.12. additionally day off)
  • Overtime is compensated by additional time off
Benefits
  • Up to 60% mobile work for suitable tasks, mostly two-person offices, bus and subway (U6) on the doorstep, free parking, pension plan of the Versorgungsanstalt des Bundes und der Länder (VBL), state-of-art work equipment, Cooperation Agreement with Sports Facility Providers for Discounted Membership
What you can expect at LRZ

Are you looking for a multifaceted and intellectually stimulating position in a dynamic, cooperative and innovative work environment? Then LRZ is the place to be for you! Here at LRZ a collegial, appreciative work environment meets an international crowd of experts who work together to advance IT services for ground-breaking research. We offer flexible work schemes for an optimal work-life balance. Our staff values their creative leeway. As an institute of the Bavarian Academy of Science and Humanities we offer all the benefits of public service. And of course, no wishes remain unfulfilled at the LRZ in terms of technical equipment. We share experiences, constantly review and improve our processes, and are proud that our service-quality and data-security are regularly certified and rated highly.

We actively promote diversity and welcome applications from talented individuals, regardless of cultural background, nationality, ethnicity, gender and sexual identity, physical abilities, religion and age. We give priority to applications from people with disabilities who are equally qualified (SGB IX).

The LRZ in a nutshell

Since 1962, Bavarian universities and research institutions have relied on the IT expertise of the Leibniz Supercomputing Centre of the Bavarian Academy of Sciences and Humanities. When it comes to the digital transformation of science, we are traditionally ahead of the game.

Here you will find information about the collection of personal data during the application process.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure
Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure

Leibniz-Rechenzentrum der Bayerischen Akademie der Wissenschaften • Garching bei München

Hybrid
EUR 60.000 - 85.000
Mobile work up to 60%
Free parking
Pension plan (VBL)
+2
System Engineer (m/w/d) Open Source Collaboration & Kubernetes
System Engineer (m/w/d) Open Source Collaboration & Kubernetes

Leibniz-Rechenzentrum • Garching bei München

Hybrid
EUR 60.000 - 75.000
JobRad
Parkplatz
VBL Altersvorsorge
+1
Wissenschaftliche Beschäftigte (m/w/d) zur Administration und Weiterentwicklung von Enterprise [...]
Wissenschaftliche Beschäftigte (m/w/d) zur Administration und Weiterentwicklung von Enterprise [...]

Leibniz-Rechenzentrum • Garching bei München

Hybrid
EUR 52.000 - 78.000
30 Tage Urlaub
Individuelle Weiterbildungen
Mobiles Arbeiten bis zu 60%
+2
System Engineer (m/w/d) Open Source Collaboration & Kubernetes
System Engineer (m/w/d) Open Source Collaboration & Kubernetes

Meyandy LLC • Garching bei München

Hybrid
EUR 70.000 - 95.000
40 Std/Woche
60% remote möglich
TV-L Entgelt
Senior AI Expert (m/f/d) for AI Computing Services
Senior AI Expert (m/f/d) for AI Computing Services

Leibniz-Rechenzentrum • Garching bei München

Hybrid
EUR 65.000
Pension plan
State-of-the-art work equipment
Free parking
Assistenz (m/w/d) für Führungskräfte im Bereich Forschung | Teilzeit 20 Std
Assistenz (m/w/d) für Führungskräfte im Bereich Forschung | Teilzeit 20 Std

Leibniz-Rechenzentrum • Garching bei München

Hybrid
EUR 27.000 - 38.000
Homeoffice bis zu 60%
Jobticket
Parkplatz
+3
Personalsachbearbeiterin (m/w/d)
Personalsachbearbeiterin (m/w/d)

Leibniz-Rechenzentrum der Bayerischen Akademie der Wissenschaften • Garching bei München

Vor Ort
EUR 60.000 - 75.000
30 Tage Urlaub
Flexible Arbeitszeiten
Mobiles Arbeiten
+2
Netzwerk-Architekt (w/m/d)
Netzwerk-Architekt (w/m/d)

Deutsches Zentrum für Luft- und Raumfahrt e. V. • Berlin

Vor Ort
EUR 70.000 - 100.000
Fort- und Weiterbildungsmöglichkeiten
Vielfalt & Gleichstellung
Bewerbungen schwerbehinderter Menschen
AI Systems and MLOps Engineer for Earth Observation
AI Systems and MLOps Engineer for Earth Observation

Euraxess • Jülich

Vor Ort
EUR 54.000 - 70.000
Comprehensive occupational health program
Flexible working options
30 days vacation
+1
Research Software Engineer in Biomedical AI for Public Health Communication (f/m/x)
Research Software Engineer in Biomedical AI for Public Health Communication (f/m/x)

Remotely • Oberschleißheim

Hybrid
EUR 70.000 - 100.000
Career Development Postdoc program
Family Support On-site kindergarten
Health Promotion
+3