Senior System Engineer (Munich, Germany)

Remotestar

München

Vor Ort

EUR 80.000 - 110.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Indefinite contract
Equal pay guaranteed
Variable performance bonus
Signing bonus
Relocation package
Private health insurance
Eligibility for educational budget
Hybrid opportunity
Flexible working hours
Career plan
Progressive Company

Zusammenfassung

Remotestar in Munich is looking for an experienced software engineer with 10+ years expertise in systems programming and deep Kubernetes knowledge. The role involves designing AI infrastructure control planes, architecting scheduling for distributed training, and optimizing networking. Ideal candidates should have a strong background with GPU ecosystems and Linux internals. The position offers an indefinite contract, private health insurance, and a hybrid working model among other benefits.

Qualifikationen

  • 10+ years of software engineering experience with strong proficiency in Python.
  • Deep understanding of Kubernetes internals and architecture.
  • Hands-on experience managing NVIDIA GPU clusters and familiarity with NVIDIA technologies.
  • Comprehensive knowledge of Linux kernel and performance tuning.
  • Mastery of declarative infrastructure tools with a focus on physical hardware.

Aufgaben

  • Designing and developing the software layer that automates lifecycle of AI infrastructure.
  • Architecting scheduling solutions for large-scale distributed training jobs.
  • Tuning the software-defined networking layer for low-latency interconnects.
  • Writing custom Kubernetes Operators and CRDs.
  • Investigating and resolving deep systems issues like PCIe bus errors.
  • Creating the 'Golden Image' for AI workloads to maximize hardware performance.

Kenntnisse

Systems Programming Expertise
Deep Kubernetes Knowledge
GPU Ecosystem Experience
Linux Internals
Infrastructure as Code
Problem Solving

Tools

Terraform
Ansible

Jobbeschreibung

About client

Well-funded and fast-growing deep-tech company founded in 2019. We are the biggest Quantum Software company in the EU. They are also one of the 100 most promising companies in AI in the world (according to CB Insights, 2023) with 150+ employees and growing, fully multicultural and international.

Requirements
  • Systems Programming Expertise: 10+ years of software engineering experience with strong proficiency in Python. You must be comfortable building system agents, APIs, and CLI tools.
  • Deep Kubernetes Knowledge: You understand K8s internals beyond simple deployment. Experience with Custom Resource Definitions (CRDs), Operators, and the Kubernetes API server architecture.
  • GPU Ecosystem Experience: Hands‑on experience managing NVIDIA GPU clusters. Familiarity with NVIDIA drivers, CUDA toolkit, and the container runtime (NVIDIA Container Toolkit).
  • Linux Internals: Deep understanding of the Linux kernel, cgroups, namespaces, and system performance tuning.
  • Infrastructure as Code: Mastery of declarative infrastructure tools (Terraform, Ansible) but with a focus on provisioning physical hardware rather than just cloud VMs.
  • Problem Solving: A proven track record of debugging complex distributed systems where the root cause could be code, network, or silicon.
Preferred qualifications
  • HPC Background: Experience working with traditional supercomputing schedulers (Slurm, PBS) or modern batch schedulers (Volcano, Kueue, Ray).
  • Bare Metal Provisioning: Experience with tools like Cluster API (CAPI), Metal3, Tinkerbell, Canonical MaaS, or OpenStack Ironic.
  • High‑Speed Networking: Knowledge of RDMA, InfiniBand, GPUDirect, and how to expose these technologies to containerized workloads.
  • AI/ML Familiarity: Understanding of how distributed training works (e.g., PyTorch Distributed, Megatron‑LM, DeepSpeed) and the infrastructure requirements of Large Language Models (LLMs).
  • Observability: Experience building monitoring for hardware health (DCGM) and distributed tracing for long‑running jobs.
Location

Applicants must have legal authorization to work in the country where the position is based.

What you will be doing
  • Building the Control Plane: Designing and developing the software layer (APIs, Controllers, Agents) that automates the lifecycle of bare‑metal AI infrastructure.
  • Orchestrating High‑Scale Compute: Architecting scheduling solutions for large‑scale distributed training jobs across massive clusters of GPUs (NVIDIA H200/B200/B300), ensuring efficient bin‑packing and gang scheduling.
  • Optimizing the Fabric: Tuning the software‑defined networking layer to support low‑latency interconnects (InfiniBand/RDMA/RoCEv2) essential for multi‑node training.
  • Developing Kubernetes Extensions: Writing custom Kubernetes Operators and CRDs to abstract complex hardware realities (topology awareness, GPU partitioning) into usable interfaces for our Data Scientists.
  • Hardware‑Level Debugging: Investigating and resolving deep systems issues, ranging from PCIe bus errors and NCCL communication timeouts to kernel panics on bare‑metal nodes.
  • Defining Standards: Creating the "Golden Image" for AI workloads, managing drivers, firmware, and OS optimizations to squeeze maximum performance out of the hardware.
Perks & Benefits
  • Indefinite contract.
  • Equal pay guaranteed.
  • Variable performance bonus.
  • Signing bonus.
  • Relocation package (if applicable).
  • Private health insurance.
  • Eligibility for educational budget according to internal policy.
  • Hybrid opportunity.
  • Flexible working hours.
  • Working in a high‑paced environment, working on cutting‑edge technologies.
  • Career plan. Opportunity to learn and teach.
  • Progressive Company. Happy people culture.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Forward Deployed Engineer
Forward Deployed Engineer

turbalance • Heidelberg

Hybrid
EUR 60.000 - 80.000
Competitive compensation
Performance-based incentives
Subsidized Deutschlandticket
+2
GPU Cluster Engineer (human)
GPU Cluster Engineer (human)

Atlas Metrics • Gemeindeverwaltungsverband Metzingen

Vor Ort
EUR 70.000 - 90.000
30 vacation days
Exclusive employee discounts
Professional development support
+1
Software Engineer - Golang / Kubernetes
Software Engineer - Golang / Kubernetes

turbalance • Heidelberg

Hybrid
EUR 70.000 - 110.000
Relocation support
Hybrid work model
High-end hardware available
+1
System Engineer (Compute Node)
System Engineer (Compute Node)

United States Digital Space LLC • Berlin

Vor Ort
EUR 85.000 - 120.000
Competitive compensation
Career growth
Flexibility
+1
Senior AI Solutions Architect - Industrial Engineering
Senior AI Solutions Architect - Industrial Engineering

NVIDIA • Deutschland

Remote
EUR 160.000 - 249.000
Equity
Benefits
Head of Compute Engineering
Head of Compute Engineering

Impossible Cloud GmbH • Hamburg

Vor Ort
EUR 80.000 - 100.000
Competitive salary
ESOP
Subsidized gym membership
+1
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Berlin

Vor Ort
EUR 120.000 - 180.000
Senior HPC Engineer, GPU Compute
Senior HPC Engineer, GPU Compute

United States Digital Space LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Competitive compensation
Career growth and learning
Mid/Senior Solution Architect - Germany
Mid/Senior Solution Architect - Germany

Multiverse Computing • München

Hybrid
EUR 70.000 - 90.000
Equal pay guaranteed
Signing bonus
Relocation package
+5
Head of Autonomy (m/f/d)
Head of Autonomy (m/f/d)

Quantum-Systems GmbH • Argelsried

Vor Ort
EUR 140.000 - 210.000
Direct impact
Flat hierarchies
30 days vacation
+2