Senior GPU and HPC Infrastructure Engineer – DGX Cloud

Jobtailor

Deutschland

Vor Ort

EUR 120.000 - 180.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Jobtailor is seeking a senior software engineer to help build and automate a platform for GPU asset provisioning, configuration, and lifecycle management across cloud providers. You will own end-to-end automation of datacenter operations for large‑scale machine learning systems, including monitoring, health management, and NVLINK topology.

You will collaborate with hardware, software, and AI training teams to ensure integration from firmware to training applications, and you will help improve

Qualifikationen

  • 10+ years of software engineering experience on large-scale production systems.
  • BS in Computer Science, Engineering, Physics, Mathematics, or comparable degree or equivalent experience.
  • Expert-level knowledge of Go and Python.
  • Expert-level Linux system administration and management.
  • Experience with cluster management systems (Kubernetes, SLURM).
  • Understanding of performance, security, and reliability in distributed systems.

Aufgaben

  • Contribute to a platform that automates GPU asset provisioning and lifecycle management across cloud providers.
  • Build end-to-end automation of datacenter operations, break/fix, and lifecycle management for large-scale Machine Learning systems.
  • Implement monitoring and health management capabilities for reliability, availability, and scalability of GPU assets.
  • Manage NVLINK topography across GPU clusters.
  • Build automated test infrastructure to qualify distributed systems for operation.
  • Collaborate with engineering teams to ensure software integration from hardware to AI training applications.

Kenntnisse

Go
Python
Linux administration

Ausbildung

BS in Computer Science
BS in Engineering
BS in Physics
BS in Mathematics

Tools

Kubernetes
SLURM

Jobbeschreibung

Responsibilities
  • Contribute to a platform that automates GPU asset provisioning, configuration, and lifecycle management across cloud providers
  • Build end-to-end automation of datacenter operations, break/fix, and lifecycle management for large‑scale Machine Learning systems
  • Implement monitoring and health management capabilities that enable reliability, availability, and scalability of GPU assets
  • Manage NVLINK topography across GPU clusters
  • Build automated test infrastructure to qualify distributed systems for operation
  • Collaborate with engineering teams to ensure software integration from hardware to AI training applications
Requirements
  • 10+ years of software engineering experience on large‑scale production systems
  • BS in Computer Science, Engineering, Physics, Mathematics, or comparable degree or equivalent experience
  • Expert-level knowledge of a systems programming language (Go, Python)
  • Expert-level knowledge of Linux system administration and management
  • Understanding of cluster management systems (Kubernetes, SLURM)
  • Understanding of performance, security, and reliability in complex distributed systems
  • Familiarity with system‑level architecture, data synchronization, fault tolerance, and state management
Core Competencies

Demonstrates expert-level knowledge in systems programming languages such as Go and Python, along with extensive experience in Linux system administration and management. Proficient in building automated solutions for large-scale Machine Learning systems and managing GPU asset provisioning and lifecycle management across cloud environments.

Certifications & Qualifications
  • BS in Computer Science
  • BS in Engineering
  • BS in Physics
  • BS in Mathematics
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Berlin

Vor Ort
EUR 120.000 - 180.000
Senior Performance Engineer
Senior Performance Engineer

NVIDIA • Deutschland

Vor Ort
EUR 120.000 - 180.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA AI • Berlin

Vor Ort
EUR 120.000 - 180.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 170.000
Senior Performance Engineer
Senior Performance Engineer

NVIDIA Corporation • Deutschland

Remote
EUR 120.000 - 180.000
Senior HPC Engineer, GPU Compute
Senior HPC Engineer, GPU Compute

Meyandy LLC • Berlin

Hybrid
EUR 120.000 - 180.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Gruppe • Berlin

Vor Ort
EUR 120.000 - 180.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Corporation • Berlin

Vor Ort
EUR 120.000 - 180.000
Head of Compute Engineering
Head of Compute Engineering

Framework Ventures • Deutschland

Hybrid
EUR 180.000 - 280.000
Equity
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Hybrid
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8