Senior GPU Cloud, K8S Expert

Jobtailor

Deutschland

Remote

EUR 90.000 - 150.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Verschicke keinen generischen Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor in Germany seeks a senior Kubernetes operations engineer to design, deploy, and operate a GPU-accelerated cloud. You will run production clusters for GPU workloads at scale and implement topology-aware scheduling.

You will integrate Slurm, Ray, and Kubeflow with Kubernetes, configure Nvidia operator and device plug‑in, enforce multi‑tenant isolation, and drive SLI/SLO metrics and incident response.

Qualifikationen

  • 5+ years of Kubernetes operations with GPU workloads.
  • Experience with Nvidia GPU Operator, device plugin, and GPU scheduling.
  • Hands-on Terraform, Helm, and GitOps workflows.
  • Strong SRE background: SLI/SLO frameworks and incident management.
  • Experience with Prometheus, Grafana, and alerting at scale.

Aufgaben

  • Design, deploy, and operate the control plane for an AI-operated GPU cloud.
  • Run production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Configure Nvidia GPU operator, device plugin, MIG, and GPU time-slicing policies.
  • Implement topology-aware scheduling using GPU locality and NVLink domain awareness.
  • Develop CRDs for GPU workload lifecycle management.
  • Integrate Slurm, Ray, and Kubeflow with Kubernetes.
  • Enforce multi-tenant isolation via namespaces, policies, RBAC, and pod security standards.
  • Automate BMaaS provisioning, tenant onboarding, lifecycle, and reclamation.
  • Build Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • Define and meet SLIs/SLOs for cluster availability, job completion, and provisioning latency.
  • Automate incident management, escalation, and post-incident reviews.
  • Operate monitoring with Prometheus, Grafana, Alertmanager, and PagerDuty.
  • Automate GPU node failure detection, drain/cordon/taint, and workload rescheduling.
  • Make the control plane safe for automated AIOps remediation.
  • Enable automated workflows for predicted GPU faults without customer impact.
  • Launch BMaaS for external tenants with self-service onboarding.
  • Publish and meet cluster availability and job-completion SLOs.

Kenntnisse

Kubernetes Operations
Nvidia GPU Operator
Terraform
Prometheus
Incident Management

Tools

Slurm
Ray
Kubeflow
Grafana
Alertmanager
PagerDuty
Helm
GitOps
ArgoCD

Jobbeschreibung

Responsibilities
  • Design, deploy, and operate the control plane for an AI-operated GPU cloud
  • Run production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs)
  • Configure Nvidia GPU operator, device plugin, MIG, and GPU time-slicing policies
  • Implement topology‑aware scheduling using GPU locality, NVLink domain awareness, and network rail affinity
  • Develop Custom Resource Definitions (CRDs) for GPU workload lifecycle management
  • Integrate Slurm, Ray, and Kubeflow with Kubernetes
  • Enforce multi‑tenant isolation through namespaces, network policies, resource quotas, RBAC, and pod security standards
  • Automate BMaaS provisioning, tenant onboarding, lifecycle, and reclamation
  • Build Terraform providers and modules for infrastructure‑as‑code across GPU clusters
  • Define and meet SLIs/SLOs for cluster availability, job completion rates, and provisioning latency
  • Automate incident management, escalation, and post‑incident reviews
  • Operate monitoring with Prometheus, Grafana, Alertmanager, and PagerDuty
  • Automate GPU node failure detection, drain/cordon/taint, and workload rescheduling
  • Make the control plane safe for automated AIOps remediation
  • Enable automated workflows for predicted GPU faults without customer impact
  • Launch BMaaS for external tenants with self‑service onboarding
  • Publish and meet cluster availability and job‑completion SLOs
Requirements
  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology‑aware scheduling and GPU‑specific resource management
  • Hands‑on experience building multi‑tenant K8S platforms with strong isolation guarantees
  • Experience with bare‑metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler
  • You've either wired an autoscaler/remediator loop into K8S or you can design one
  • Runbook‑as‑code mindset — every SRE playbook you write should be executable by the platform
Core Competencies

Demonstrates expertise in Kubernetes operations, particularly in managing GPU workloads and implementing topology‑aware scheduling. Proficient in automation, incident management, and building multi‑tenant platforms with strong isolation guarantees.

Highest‑signal resume keywords
  • Kubernetes Operations
  • Nvidia GPU Operator
  • Terraform
  • Prometheus
  • Incident Management
ATS Optimization Keywords
Hard Skills
  • GPU Workload Management
  • Topology‑Aware Scheduling
  • Custom Resource Definitions (CRDs)
  • Multi‑Tenant Kubernetes Platforms
  • Bare‑Metal Server Provisioning
  • Programming in Go
  • Programming in Python
  • AIOps
  • SLI/SLO Frameworks
  • Runbook‑as‑Code
Industry Keywords
  • BMaaS
  • Resource Quotas
  • RBAC
  • Pod Security Standards
  • Network Policies
  • Capacity Planning
  • Incident Management Automation
  • Workload Rescheduling
  • GPU Time‑Slicing Policies
  • Network Rail Affinity
Tools & Technologies
  • Kubernetes
  • Slurm
  • Ray
  • Kubeflow
  • Grafana
  • Alertmanager
  • PagerDuty
  • Helm
  • GitOps
  • ArgoCD
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

HPC Infrastructure Engineer – GPU Clusters
HPC Infrastructure Engineer – GPU Clusters

Jobtailor • Deutschland

Hybrid
EUR 80.000 - 140.000
Technical Lead – GPU Infrastructure
Technical Lead – GPU Infrastructure

Jobtailor • Deutschland

Remote
EUR 120.000 - 160.000
SRE Monitoring Platform Software Engineer, Entry Level
SRE Monitoring Platform Software Engineer, Entry Level

Jobtailor • Deutschland

Remote
EUR 42.000 - 64.000
Compute Solution Architect
Compute Solution Architect

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Senior GPU Cloud Storage Solutions Expert – SRE SME
Senior GPU Cloud Storage Solutions Expert – SRE SME

Jobtailor • Deutschland

Remote
EUR 90.000 - 140.000
Software Platform Support Engineer – GPU Cloud
Software Platform Support Engineer – GPU Cloud

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 140.000
Senior System Software Engineer, Software Defined Networking
Senior System Software Engineer, Software Defined Networking

Jobtailor • Deutschland

Remote
EUR 90.000 - 140.000
DevOps Engineer – On-Premises Cloud
DevOps Engineer – On-Premises Cloud

Jobtailor • Bonn

Vor Ort
EUR 90.000 - 130.000
Senior Technical Support Engineer – Slurm
Senior Technical Support Engineer – Slurm

Jobtailor • Deutschland

Remote
EUR 70.000 - 100.000