GPU Cluster Engineer - Scalable AI Infrastructure

NEURA Robotics

Germany (OH)

On-site

USD 140,000 - 195,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NEURA Robotics seeks a senior infrastructure engineer to own and evolve its large-scale GPU cluster platform on AWS HyperPod. You will design, operate, and optimize the infra that enables foundation model training and customer workloads.

You will implement Slurm and Kubernetes orchestration, build self-service tooling, automate reliability, and partner with ML researchers, product teams, and AWS to improve throughput, stability, and cost awareness across teams.

Qualifications

  • 5+ years of experience in infrastructure or systems engineering.
  • Strong experience with GPU cluster or HPC operations.
  • Hands-on with AWS HyperPod and GPU instances.
  • Solid understanding of Slurm and Kubernetes.
  • Experience building self-service tooling and end-user docs.

Responsibilities

  • Own the design, deployment, and operation of NEURA's GPU cluster platform on AWS HyperPod.
  • Develop and refine orchestration for HyperPod/Slurm and HyperPod/EKS.
  • Improve cluster stability: failure detection, automatic recovery, and fault-tolerant multi-node training.
  • Build workload prioritization to fairly share cluster capacity across teams.
  • Optimize end-to-end GPU utilization across compute, memory, networking, and storage.
  • Collaborate with AWS HyperPod teams and internal orgs to influence roadmap.
  • Deliver self-service tooling for researchers to launch and monitor training jobs.

Skills

GPU cluster operations
AWS HyperPod
Slurm
Kubernetes
Distributed training
Self-service tooling
Infrastructure documentation
Cost management
Cross-functional collaboration
English communication
German knowledge

Tools

Terraform

Job description

NEURA Robotics seeks a senior infrastructure engineer to own and evolve its large-scale GPU cluster platform on AWS HyperPod. You will design, operate, and optimize the infra that enables foundation model training and customer workloads.

You will implement Slurm and Kubernetes orchestration, build self-service tooling, automate reliability, and partner with ML researchers, product teams, and AWS to improve throughput, stability, and cost awareness across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer (human)
GPU Cluster Engineer (human)

NEURA Robotics • Germany (OH)

On-site
USD 140,000 - 195,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Cluster Infra Engineer - Reliability & Automation
GPU Cluster Infra Engineer - Reliability & Automation

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Health benefits
401k matching
+2