Technical Lead – GPU Infrastructure

Jobtailor

Deutschland

Remote

EUR 120.000 - 160.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor is seeking a Senior Platform Architect to own the end-to-end platform architecture, lead a distributed team across backend, frontend, DevOps, QA, and documentation, and establish standards and release gates for a research-oriented HPC environment.

You will design, build, and operate a managed Slurm service, manage GPU fleets, and drive reliability with SLOs, incident response, and on-call excellence.

Qualifikationen

  • Eight or more years of hands-on engineering experience.
  • Led teams that build and operate infrastructure platforms.
  • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
  • Hands-on Slurm experience at scale including slurmctl, slurmdbd and QoS.

Aufgaben

  • Own end-to-end platform architecture and design reviews.
  • Lead a distributed team across backend, frontend, DevOps, QA, and documentation.
  • Establish engineering standards, release gates, and capacity planning.
  • Operate a managed Slurm service for research users.
  • Manage Kubernetes cluster lifecycle on bare metal and NVIDIA GPU integration.

Kenntnisse

Leadership
Slurm at scale
Kubernetes
Linux systems
GPU management
Prometheus/Grafana
NVIDIA GPUs
CI/CD/Automation

Ausbildung

Bachelor's or Master's in CS/Engineering

Tools

Lustre
NFS
NVSwitch
VFIO

Jobbeschreibung

  • Own the end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintenance of the baseline
  • Lead and line-manage a distributed team of about twelve engineers across backend, frontend, DevOps, QA, and documentation
  • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input
  • Design, build, and operate a managed Slurm service for research users
  • Own controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation
  • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal
  • Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement
  • Own managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
  • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers
  • Lead incident response, post-incident reviews, and development of a sustainable on-call model
  • Serve as the primary technical interface to infrastructure partners and vendors
  • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing
  • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity
  • Complete the platform team and set the technical bar for new engineers
Requirements
  • Eight or more years of hands-on engineering experience
  • At least three years leading teams that build and operate infrastructure platforms other teams depend on
  • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
  • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system
  • Experience operating an HPC or GPU training cluster for a research population is ideally preferred
  • Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance
  • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems
  • Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning
  • Production Kubernetes operations experience, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design
  • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets
  • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review
  • Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions
  • Experience shipping a platform with real users, such as a multi-tenant IaaS/PaaS or research computing service
  • People management across time zones, cross-track review, written architecture decisions, and partner/executive communication
  • Excellent written and spoken English
  • Based between UTC and UTC+5:30
  • Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers
  • Desirable: modern serving stacks such as vLLM, SGLang, and TensorRT-LLM
  • Desirable: VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and hardware-provider partnership experience
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

HPC Infrastructure Engineer – GPU Clusters
HPC Infrastructure Engineer – GPU Clusters

Jobtailor • Deutschland

Hybrid
EUR 80.000 - 140.000
Senior GPU Cloud, K8S Expert
Senior GPU Cloud, K8S Expert

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 140.000
HPC Cluster Architect
HPC Cluster Architect

nexgencloud • Deutschland

Hybrid
EUR 120.000 - 180.000
Annual discretionary bonus
25 days holiday
Remote or hybrid options
+1
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Gruppe • Berlin

Vor Ort
EUR 120.000 - 180.000
Software Platform Support Engineer – GPU Cloud
Software Platform Support Engineer – GPU Cloud

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Compute Solution Architect
Compute Solution Architect

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Senior HPC GPU Cluster Lead for Deep Learning Infra
Senior HPC GPU Cluster Lead for Deep Learning Infra

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Corporation • Berlin

Vor Ort
EUR 120.000 - 180.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 170.000