GPU Cluster Engineer (human)

NEURA Robotics

Germany (OH)

On-site

USD 140,000 - 195,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NEURA Robotics seeks a senior infrastructure engineer to own and evolve its large-scale GPU cluster platform on AWS HyperPod. You will design, operate, and optimize the infra that enables foundation model training and customer workloads.

You will implement Slurm and Kubernetes orchestration, build self-service tooling, automate reliability, and partner with ML researchers, product teams, and AWS to improve throughput, stability, and cost awareness across teams.

Qualifications

  • 5+ years of experience in infrastructure or systems engineering.
  • Strong experience with GPU cluster or HPC operations.
  • Hands-on with AWS HyperPod and GPU instances.
  • Solid understanding of Slurm and Kubernetes.
  • Experience building self-service tooling and end-user docs.

Responsibilities

  • Own the design, deployment, and operation of NEURA's GPU cluster platform on AWS HyperPod.
  • Develop and refine orchestration for HyperPod/Slurm and HyperPod/EKS.
  • Improve cluster stability: failure detection, automatic recovery, and fault-tolerant multi-node training.
  • Build workload prioritization to fairly share cluster capacity across teams.
  • Optimize end-to-end GPU utilization across compute, memory, networking, and storage.
  • Collaborate with AWS HyperPod teams and internal orgs to influence roadmap.
  • Deliver self-service tooling for researchers to launch and monitor training jobs.

Skills

GPU cluster operations
AWS HyperPod
Slurm
Kubernetes
Distributed training
Self-service tooling
Infrastructure documentation
Cost management
Cross-functional collaboration
English communication
German knowledge

Tools

Terraform

Job description

Your mission & challenges
  • You are the go-to expert for NEURA's GPU cluster infrastructure - a large-scale AWS HyperPod environment running cutting-edge GPU instances for foundation model training and customer fine-tuning workloads. You design the operational framework, build self-service tooling for ML teams, and work directly with AWS to influence the platform at the hyperscaler level.

  • Your focus is on cluster engineering and operations — not on ML research itself, but on making sure the people doing that research have rock-solid, efficient, and accessible infrastructure under them.

  • Setting up, configuring, and continuously evolving NEURA's HyperPod clusters, including HyperPod/Slurm and HyperPod/EKS orchestration models.

  • Designing and implementing strategies for cluster stability: node failure detection, automated job recovery, checkpoint coordination, and fault-tolerant multi-node training workflows.

  • Providing a workload priority management framework that allows multiple teams and use cases like foundation model pretraining, fine-tuning, customer workloads, to share cluster capacity efficiently and fairly.

  • Optimizing end-to-end GPU utilization: identifying and resolving bottlenecks across compute, GPU memory, EFA networking, and storage throughput.

  • Working directly and closely with the AWS HyperPod product and solutions engineering teams, escalating operational issues, sharing learnings from one of the platform's largest deployments, and placing concrete requirements on the roadmap.

  • Providing self-service tooling that allows ML researchers and engineers to launch, monitor, and manage training jobs independently, without requiring infrastructure intervention for routine operations.

  • Developing onboarding documentation, training materials, and internal workshops that enable users to operate efficiently, follow best practices, and understand cost implications of their workloads.

  • Infrastructure as Code is a given for you. Every cluster configuration, every operational change, every new environment is code first.

  • Owning the cost and capacity strategy: Spot instance management, Reserved Instance planning, Savings Plans, and ongoing commitment negotiations with AWS.

What we can look forward to
  • 5+ years of experience in infrastructure or systems engineering, with a strong focus on GPU cluster or HPC operations.

  • Deep hands-on experience with AWS HyperPod and AWS instances; direct prior experience with HyperPod is a strong differentiator.

  • Solid understanding of both Slurm and Kubernetes as cluster orchestration layers, and the ability to evaluate their trade-offs for large-scale GPU workloads.

  • Practical knowledge of distributed training - you understand what affects throughput and how to debug it.

  • Experience building self-service tooling and operational documentation for technical end users.

  • You make complex infrastructure accessible, not just functional.

  • Strong understanding of cloud cost management at scale: Spot interruption handling, capacity reservations, cost attribution across teams and workloads.

  • Comfort working across organizational boundaries — your primary partners are ML researchers, but you'll also work closely with product, finance, and cloud vendor teams.

  • Strong English communication skills. German is a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU Cluster Engineer - Scalable AI Infrastructure
GPU Cluster Engineer - Scalable AI Infrastructure

NEURA Robotics • Germany (OH)

On-site
USD 140,000 - 195,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Germany (OH)

On-site
USD 58,000 - 102,000