GPU Cluster Engineer - Scalable AI Infrastructure

NEURA Robotics

Germany (OH)

On-site

USD 140,000 - 195,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NEURA Robotics seeks a senior infrastructure engineer to own and evolve its large-scale GPU cluster platform on AWS HyperPod. You will design, operate, and optimize the infra that enables foundation model training and customer workloads.

You will implement Slurm and Kubernetes orchestration, build self-service tooling, automate reliability, and partner with ML researchers, product teams, and AWS to improve throughput, stability, and cost awareness across teams.

Qualifications

  • 5+ years of experience in infrastructure or systems engineering.
  • Strong experience with GPU cluster or HPC operations.
  • Hands-on with AWS HyperPod and GPU instances.
  • Solid understanding of Slurm and Kubernetes.
  • Experience building self-service tooling and end-user docs.

Responsibilities

  • Own the design, deployment, and operation of NEURA's GPU cluster platform on AWS HyperPod.
  • Develop and refine orchestration for HyperPod/Slurm and HyperPod/EKS.
  • Improve cluster stability: failure detection, automatic recovery, and fault-tolerant multi-node training.
  • Build workload prioritization to fairly share cluster capacity across teams.
  • Optimize end-to-end GPU utilization across compute, memory, networking, and storage.
  • Collaborate with AWS HyperPod teams and internal orgs to influence roadmap.
  • Deliver self-service tooling for researchers to launch and monitor training jobs.

Skills

GPU cluster operations
AWS HyperPod
Slurm
Kubernetes
Distributed training
Self-service tooling
Infrastructure documentation
Cost management
Cross-functional collaboration
English communication
German knowledge

Tools

Terraform

Job description

NEURA Robotics seeks a senior infrastructure engineer to own and evolve its large-scale GPU cluster platform on AWS HyperPod. You will design, operate, and optimize the infra that enables foundation model training and customer workloads.

You will implement Slurm and Kubernetes orchestration, build self-service tooling, automate reliability, and partner with ML researchers, product teams, and AWS to improve throughput, stability, and cost awareness across teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Equity incentives
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Platform Architect - HPC, Kubernetes
Platform Architect - HPC, Kubernetes

EPAM Systems Inc • United States

Remote
USD 140,000 - 230,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff, AI Compute & Data Infrastructure Vinci
Member of Technical Staff, AI Compute & Data Infrastructure Vinci

CDFAM - Computational Design Symposium • Palo Alto (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000