Senior System Engineer (Munich, Germany)

remotestar-team

Cambourne

Hybrid

GBP 90,000 - 150,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Indefinite contract
Equal pay guaranteed
Variable performance bonus
Signing bonus
Relocation package
Private health insurance
Educational budget
Hybrid opportunity
Flexible working hours
High-paced cutting-edge tech

Job summary

Unknown company is seeking a senior systems engineer to design and build the control plane for bare-metal AI infrastructure. You will architect GPU scheduling across massive clusters, optimize networking for low latency, and extend Kubernetes with Operators and CRDs to expose hardware realities to data scientists.

Deep kernel and device driver work is required. You will debug complex issues from PCIe to NCCL timeouts, and define standards for drivers, firmware, and OS optimizations to maximize

Qualifications

  • 10+ years of software engineering experience with Python.
  • Deep Kubernetes knowledge: CRDs, Operators, API server architecture.
  • Hands-on experience managing NVIDIA GPU clusters, drivers, CUDA, and NVIDIA Container Toolkit.
  • Deep Linux kernel, cgroups, namespaces, and performance tuning.
  • Terraform/Ansible for provisioning physical hardware.
  • Experience debugging distributed systems across code, network, or silicon.

Responsibilities

  • Build the Control Plane to automate bare-metal AI infrastructure lifecycles.
  • Architect scheduling for large-scale GPU clusters with efficient bin-packing and gang scheduling.
  • Tune the software-defined networking layer for low-latency interconnects (InfiniBand/RDMA).
  • Develop Kubernetes Extensions via Operators and CRDs for hardware abstractions.
  • Debug PCIe, NCCL timeouts, and kernel panics on bare-metal nodes.
  • Define the Golden Image for AI workloads including drivers and firmware.

Skills

Python
Kubernetes
Linux Internals
GPU Clusters
Infrastructure as Code
Troubleshooting

Tools

NVIDIA Container Toolkit
Cluster API (CAPI)
Metal3
Tinkerbell
Canonical MaaS
OpenStack Ironic

Job description

About client :

Well-funded and fast-growing deep-tech company founded in 2019. We are the biggest Quantum Software company in the EU. They are also one of the 100 most promising companies in AI in the world (according to CB Insights, 2023) with 150+ employees and growing, fully multicultural and international.

Requirements
  • Systems Programming Expertise: 10+ years of software engineering experience with strong proficiency in Python. You must be comfortable building system agents, APIs, and CLI tools.
  • Deep Kubernetes Knowledge: You understand K8s internals beyond simple deployment. Experience with Custom Resource Definitions (CRDs), Operators, and the Kubernetes API server architecture.
  • GPU Ecosystem Experience: Hands-on experience managing NVIDIA GPU clusters. Familiarity with NVIDIA drivers, CUDA toolkit, and the container runtime (NVIDIA Container Toolkit).
  • Linux Internals: Deep understanding of the Linux kernel, cgroups, namespaces, and system performance tuning.
  • Infrastructure as Code: Mastery of declarative infrastructure tools (Terraform, Ansible) but with a focus on provisioning physical hardware rather than just cloud VMs.
  • Problem Solving: A proven track record of debugging complex distributed systems where the root cause could be code, network, or silicon.
Preferred qualifications
  • HPC Background: Experience working with traditional supercomputing schedulers (Slurm, PBS) or modern batch schedulers (Volcano, Kueue, Ray).
  • Bare Metal Provisioning: Experience with tools like Cluster API (CAPI), Metal3, Tinkerbell, Canonical MaaS, or OpenStack Ironic.
  • High-Speed Networking: Knowledge of RDMA, InfiniBand, GPUDirect, and how to expose these technologies to containerized workloads.
  • AI/ML Familiarity: Understanding of how distributed training works (e.g., PyTorch Distributed, Megatron-LM, DeepSpeed) and the infrastructure requirements of Large Language Models (LLMs).
  • Observability: Experience building monitoring for hardware health (DCGM) and distributed tracing for long-running jobs.
Location: Applicants must have legal authorization to work in the country where the position is based
What you will be doing
  • Building the Control Plane: Designing and developing the software layer (APIs, Controllers, Agents) that automates the lifecycle of bare-metal AI infrastructure.
  • Orchestrating High-Scale Compute: Architecting scheduling solutions for large-scale distributed training jobs across massive clusters of GPUs (NVIDIA H200/B200/B300), ensuring efficient bin-packing and gang scheduling.
  • Optimizing the Fabric: Tuning the software-defined networking layer to support low-latency interconnects (InfiniBand/RDMA/RoCEv2) essential for multi-node training.
  • Developing Kubernetes Extensions: Writing custom Kubernetes Operators and CRDs to abstract complex hardware realities (topology awareness, GPU partitioning) into usable interfaces for our Data Scientists.
  • Hardware-Level Debugging: Investigating and resolving deep systems issues, ranging from PCIe bus errors and NCCL communication timeouts to kernel panics on bare-metal nodes.
  • Defining Standards: Creating the "Golden Image" for AI workloads, managing drivers, firmware, and OS optimizations to squeeze maximum performance out of the hardware.
Perks & Benefits
  • Indefinite contract.
  • Equal pay guaranteed.
  • Variable performance bonus.
  • Signing bonus.
  • Relocation package (if applicable).
  • Private health insurance.
  • Eligibility for educational budget according to internal policy.
  • Hybrid opportunity.
  • Flexible working hours.
  • Working in a high paced environment, working on cutting edge technologies.
  • Career plan. Opportunity to learn and teach.
  • Progressive Company. Happy people culture
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform Engineer (Product Initiatives) - Systems Integrator
Senior Platform Engineer (Product Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Remote
GBP 120,000 - 190,000
High-Upside Equity
Flexible remote setup
Work-Life Balance
+1
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Network Engineer
Network Engineer

asobbi • United Kingdom

Remote
GBP 53,000 - 69,000
Highly competitive package with equity
Dynamic progression plan
Human-first flexibility
Senior DevOps Engineer
Senior DevOps Engineer

Quantum Motion Technologies • Greater London

Hybrid
GBP 90,000 - 150,000
Health insurance
Share options scheme
Private Medical Insurance
+1
IT Lead Engineer
IT Lead Engineer

PhysicsX • Greater London

On-site
GBP 90,000 - 140,000
Equity options
10% employer pension contribution
Free office lunches
+10
Mid/Senior Solution Architect - UK
Mid/Senior Solution Architect - UK

Multiverse Computing • Greater London

Hybrid
GBP 90,000 - 140,000
Hybrid opportunity
Signing bonus
Relocation package
+4
Mid/Senior Solution Architect - UK
Mid/Senior Solution Architect - UK

Multiverse Computing • City Of London

Hybrid
GBP 90,000 - 130,000
Equal pay guaranteed
Signing bonus
Relocation package
+5
Deployment Engineering Director, Systems Engineering New UK
Deployment Engineering Director, Systems Engineering New UK

Nscale • United Kingdom

On-site
GBP 120,000 - 180,000
Senior Software Engineer, Infrastructure - Python & Kubernetes
Senior Software Engineer, Infrastructure - Python & Kubernetes

PhysicsX • Greater London

Hybrid
GBP 90,000 - 120,000
Equity options
10% employer pension contribution
Free office lunches
+2
Senior HPC Engineer, GPU Compute
Senior HPC Engineer, GPU Compute

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive compensation
Career growth
Flexibility and ownership
+3