Senior GPU Cloud Storage Solutions Expert – SRE SME

Jobtailor

Deutschland

Remote

EUR 90.000 - 140.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor seeks a seasoned Storage Operations Engineer to deploy and manage parallel storage systems (WEKA, VAST Data, Ceph, DDN/Lustre) and optimize AI workload I/O patterns in a large-scale environment.

You will implement multi-tenant QoS, configure high-speed networking (NFS over RDMA, NVMe-oF), and build automation from runbooks to code. Strong Linux, benchmarking (fio/IOR/mdtest) and disaster recovery planning are essential.

Qualifikationen

  • 5+ years in enterprise or HPC storage operations, with at least 2 years supporting AI/ML workloads.
  • Hands-on deployment and operations experience with WEKA, VAST Data, Ceph, or DDN/Lustre.
  • Strong understanding of AI training I/O patterns and high-performance storage networking.
  • Experience shipping anomaly detectors for storage telemetry and multi-tenant isolation.

Aufgaben

  • Deploy and operate parallel/distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre.
  • Design architectures optimized for AI workloads with efficient I/O patterns.
  • Implement multi-tenant storage isolation with QoS, quotas, and access controls.
  • Configure and optimize GPU Direct Storage and RDMA-based data transfer.
  • Deploy and manage storage networking (NFS over RDMA, NVMe-oF).
  • Profile and tune storage performance using fio, IOR, and mdtest.
  • Plan capacity aligned with GPU cluster growth and workload projections.

Kenntnisse

Storage operations
AI/ML workloads
Linux systems
Performance tuning
Multi-tenant isolation
Storage networking
Telemetry instrumentation
Automation mindset

Tools

WEKA
VAST Data
Ceph
DDN/Lustre

Jobbeschreibung

  • Deploy and operate parallel/distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre
  • Design storage architectures optimized for AI workload patterns such as checkpoint I/O bursts, sequential dataset reads, and KV cache for inference
  • Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls
  • Configure and optimize GPU Direct Storage for direct GPU-to-storage data paths
  • Deploy and manage storage networking including NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX for cluster-wide storage orchestration
  • Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest
  • Own the runbook for common failure modes
  • Plan storage capacity aligned with GPU cluster growth and customer workload projections
  • Manage firmware, data migration, and disaster recovery procedures
  • Instrument storage telemetry including IO tail latency, checkpoint durations, NVMe SMART, filesystem health, and RDMA counters
  • Feed telemetry into the platform team's metrics, logs, and traces store
  • Partner with the platform team to define the storage-fault predictor, including signals, labels, and false-positive tolerances
  • Convert novel incidents into automation, progressing from SOPs to runbook-as-code and agent-executable remediation
  • Deliver observability and a baseline predictor for the top three storage-fault classes
  • Reduce storage-incident MTTR
  • Design storage for Nvidia GB200-class clusters
Requirements
  • 5+ years in enterprise or HPC storage operations, with at least 2 years supporting AI/ML workloads
  • Hands-on deployment and operations experience with at least two of: WEKA, VAST Data, Ceph, DDN/Lustre
  • Strong understanding of AI training I/O patterns: checkpoint frequency, dataset loading, shuffle buffers
  • Experience with high-performance storage networking (NFS over RDMA, NVMe-oF)
  • Knowledge of GPU Direct Storage and RDMA-based data transfer
  • Proficiency in storage performance benchmarking and tuning (fio, IOR, mdtest)
  • Experience implementing multi-tenant storage with isolation and QoS
  • Strong Linux systems knowledge (kernel tuning, filesystem internals, block device management)
  • Experience shipping an anomaly detector for storage/IO telemetry or ability to articulate the labels and features needed
  • Runbook-as-code mindset, with every SOP executable by a machine within a quarter
Core Competencies

Demonstrates expertise in deploying and managing parallel and distributed storage systems optimized for AI workloads, with a strong focus on performance tuning, multi-tenant isolation, and telemetry instrumentation. Proficient in high-performance storage networking and capable of implementing automation for storage operations.

Highest-signal resume keywords
  • WEKA Deployment
  • Ceph Operations
  • GPU Direct Storage
  • Storage Performance Benchmarking
  • Multi-Tenant Storage Isolation
ATS Optimization Keywords
Hard Skills
  • Storage Architecture Design
  • AI Workload Optimization
  • Storage Performance Tuning
  • Linux Systems Knowledge
  • Anomaly Detection for Telemetry
Soft Skills
  • Problem-Solving
  • Collaboration
Industry Keywords
  • Enterprise Storage Operations
  • HPC Storage
  • AI/ML Workloads
  • Telemetry Instrumentation
  • Disaster Recovery Procedures
Tools & Technologies
  • NFS over RDMA
  • NVMe-oF
  • Fio
  • IOR
  • Mdtest
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Compute Solution Architect
Compute Solution Architect

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
HPC Infrastructure Engineer – GPU Clusters
HPC Infrastructure Engineer – GPU Clusters

Jobtailor • Deutschland

Hybrid
EUR 80.000 - 140.000
Senior GPU Cloud, K8S Expert
Senior GPU Cloud, K8S Expert

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Senior Solutions Architect, Physical AI Cloud
Senior Solutions Architect, Physical AI Cloud

Jobtailor • Deutschland

Vor Ort
EUR 110.000 - 170.000
Technical Lead – GPU Infrastructure
Technical Lead – GPU Infrastructure

Jobtailor • Deutschland

Remote
EUR 120.000 - 160.000
Software Platform Support Engineer – GPU Cloud
Software Platform Support Engineer – GPU Cloud

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Advisory Sales Engineer - Strategic AI
Advisory Sales Engineer - Strategic AI

DDN • Berlin

Vor Ort
EUR 80.000 - 110.000
Senior Compute Network Engineer
Senior Compute Network Engineer

Jobtailor • Deutschland

Remote
EUR 90.000 - 120.000
NetApp Administrator
NetApp Administrator

Jobtailor • Bremen

Vor Ort
EUR 55.000 - 85.000
Senior System Software Engineer, Software Defined Networking
Senior System Software Engineer, Software Defined Networking

Jobtailor • Deutschland

Remote
EUR 90.000 - 140.000