HPC Infrastructure Engineer – GPU Clusters

Jobtailor

Deutschland

Hybrid

EUR 80.000 - 140.000

Vollzeit

Vor 12 Tagen
Bewerbungsgenerator

Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor is seeking a hands-on GPU infrastructure engineer to manage and optimize our large-scale GPU compute environment. You will provision, monitor, and upgrade hardware, build automation for health checks and remediation, and own the stack beneath ML training code.

You will run Slurm scheduling, design high-speed storage, and work with researchers to boost training throughput, while maintaining security and efficiency. Occasional datacenter work may be required.

Qualifikationen

  • Production experience running large-scale Linux server or GPU environments.
  • Strong knowledge of NVIDIA drivers, CUDA, NCCL, and DCGM, or deep systems experience with ability to learn hardware stacks quickly.
  • Experience with bare-metal environments, server hardware, and high-speed networking.
  • Proficiency in Python and/or Bash automation.
  • Experience with infrastructure-as-code tools such as Ansible or Terraform.
  • Ability to analyze metrics, logs, and PromQL.
  • Experience supporting ML training workloads from the infrastructure side (nice to have).
  • Experience evaluating and working with GPU cloud providers (nice to have).
  • Experience with parallel filesystems such as WEKA or VAST, or large-scale object storage (nice to have).
  • Experience with BMC/IPMI/Redfish automation and PXE provisioning at scale (nice to have).
  • Power and cooling awareness for dense GPU deployments (nice to have).
  • Willingness to perform datacenter trips and hands-on hardware work.

Aufgaben

  • Operate and improve the GPU fleet end to end, including provisioning, scheduling, monitoring, upgrades, and capacity planning.
  • Build automation for node health checks, automated draining and remediation, and burn-in pipelines.
  • Own the infrastructure stack beneath training code, including OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and InfiniBand/RoCE networking.
  • Run and tune job scheduling with Slurm or similar systems.
  • Build and maintain high-performance storage for datasets and checkpoints.
  • Investigate and resolve performance problems involving stragglers, degraded links, thermal issues, and faulty GPUs.
  • Evaluate rented GPU capacity by benchmarking providers, validating capacity, and enforcing SLAs.
  • Perform hands-on hardware work, including racking, cabling, and diagnostics.
  • Coordinate with datacenter staff and vendors.
  • Maintain cluster security through access control, network isolation, and secrets management.
  • Work directly with researchers to improve training throughput and researcher velocity.

Kenntnisse

GPU infra mgmt
NVIDIA drivers & CUDA
Python scripting
Bash automation
IaC: Ansible
IaC: Terraform
Metrics & logs
ML training infra
GPU cloud providers
WEKA/VAST
BMC/IPMI/Redfish
PXE provisioning
Power & cooling
Datacenter trips
Slurm scheduling

Tools

NCCL
DCGM
PromQL

Jobbeschreibung

  • Operate and improve the GPU fleet end to end, including provisioning, scheduling, monitoring, upgrades, and capacity planning
  • Build automation for node health checks, automated draining and remediation, and burn-in pipelines
  • Own the infrastructure stack beneath training code, including OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and InfiniBand/RoCE networking
  • Run and tune job scheduling with Slurm or similar systems
  • Build and maintain high-performance storage for datasets and checkpoints
  • Investigate and resolve performance problems involving stragglers, degraded links, thermal issues, and faulty GPUs
  • Evaluate rented GPU capacity by benchmarking providers, validating capacity, and enforcing SLAs
  • Perform hands-on hardware work, including racking, cabling, and diagnostics
  • Coordinate with datacenter staff and vendors
  • Maintain cluster security through access control, network isolation, and secrets management
  • Work directly with researchers to improve training throughput and researcher velocity

Requirements

  • Production experience running large-scale Linux server or GPU environments
  • Strong knowledge of NVIDIA drivers, CUDA, NCCL, and DCGM, or deep systems experience with ability to learn hardware stacks quickly
  • Experience with bare-metal environments, server hardware, and high-speed networking
  • Proficiency in Python and/or Bash automation
  • Experience with infrastructure-as-code tools such as Ansible or Terraform
  • Ability to analyze metrics, logs, and PromQL
  • Experience supporting ML training workloads from the infrastructure side (nice to have)
  • Experience evaluating and working with GPU cloud providers (nice to have)
  • Experience with parallel filesystems such as WEKA or VAST, or large-scale object storage (nice to have)
  • Experience with BMC/IPMI/Redfish automation and PXE provisioning at scale (nice to have)
  • Power and cooling awareness for dense GPU deployments (nice to have)
  • Willingness to perform datacenter trips and hands-on hardware work

Core Competencies

Demonstrates expertise in managing and optimizing GPU infrastructure, including proficiency in NVIDIA drivers, CUDA, and automation tools. Capable of performing hands-on hardware work while ensuring cluster security and high-performance storage management.

Highest-signal resume keywords

  • GPU Infrastructure Management
  • NVIDIA Drivers and CUDA Proficiency
  • Python and Bash Automation
  • Infrastructure-as-Code Tools (Ansible, Terraform)
  • Performance Analysis and Troubleshooting

ATS Optimization Keywords

Hard Skills

  • Linux Server Management
  • GPU Environment Operations
  • Job Scheduling with Slurm
  • High-Speed Networking
  • Metrics and Log Analysis
  • Bare-Metal Environments
  • Parallel Filesystems (WEKA, VAST)
  • BMC/IPMI/Redfish Automation
  • PXE Provisioning
  • Capacity Planning

Soft Skills

  • Collaboration with Researchers
  • Problem-Solving
  • Communication with Datacenter Staff

Industry Keywords

  • GPU Cloud Providers
  • ML Training Workloads
  • Cluster Security
  • Access Control
  • Network Isolation

Tools & Technologies

  • NCCL
  • DCGM
  • PromQL
  • InfiniBand/RoCE Networking
  • Automated Health Checks
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior GPU Cloud, K8S Expert
Senior GPU Cloud, K8S Expert

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Technical Lead – GPU Infrastructure
Technical Lead – GPU Infrastructure

Jobtailor • Deutschland

Remote
EUR 120.000 - 160.000
Compute Solution Architect
Compute Solution Architect

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Senior GPU Cloud Storage Solutions Expert – SRE SME
Senior GPU Cloud Storage Solutions Expert – SRE SME

Jobtailor • Deutschland

Remote
EUR 90.000 - 140.000
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 140.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 170.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Gruppe • Berlin

Vor Ort
EUR 120.000 - 180.000
HPC Cluster Architect
HPC Cluster Architect

nexgencloud • Deutschland

Hybrid
EUR 120.000 - 180.000
Annual discretionary bonus
25 days holiday
Remote or hybrid options
+1
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Corporation • Berlin

Vor Ort
EUR 120.000 - 180.000