GPU Cluster Engineer (human)

Atlas Metrics

Gemeindeverwaltungsverband Metzingen

Vor Ort

EUR 70.000 - 90.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

30 vacation days
Exclusive employee discounts
Professional development support
Flexible work hours

Zusammenfassung

Atlas Metrics is looking for a GPU Cluster Engineer to manage and enhance our cutting-edge AWS HyperPod infrastructure. In this full-time role, you'll ensure efficient and accessible resources for ML teams while shaping the future of human-robot collaboration. Ideal candidates will have over 5 years of experience in infrastructure engineering, especially in GPU operations.

The position offers a dynamic work culture with flexible hours, professional growth opportunities, and a competitive salary. Join us to redefine the landscape of technology!

Qualifikationen

  • 5+ years of experience in infrastructure or systems engineering focused on GPU cluster operations.
  • Deep hands-on experience with AWS HyperPod and AWS instances.
  • Solid understanding of Slurm and Kubernetes as orchestration layers.
  • Experience building self-service tooling for technical end users.
  • Strong understanding of cloud cost management at scale.

Aufgaben

  • Manage NEURA's GPU cluster infrastructure for efficient model training.
  • Design operational frameworks and self-service tooling for ML teams.
  • Continuously evolve HyperPod clusters to optimize performance.
  • Implement strategies for cluster stability and automated job recovery.
  • Resolve bottlenecks in performance across software and hardware.

Kenntnisse

Infrastructure management
AWS HyperPod
Cluster orchestration (Slurm, Kubernetes)
Distributed training knowledge
Cloud cost management
Self-service tooling development
Strong English communication
German language knowledge

Jobbeschreibung

GPU Cluster Engineer (human)

Neura Robotics • Metzingen

Full-time

from today

Shape the Future of Human-Robot Collaboration

In the Software Department, you're shaping robotic solutions that redefine human-machine collaboration. You'll work with cutting-edge technology, setting industry-changing standards. Not only will you help develop our solutions, but you'll also set new trends and drive innovations forward. In an agile and interdisciplinary team, you'll engage in exciting projects. With clear Scrum processes like daily stand‑ups, sprint planning, and reviews, you remain flexible and efficient. Collaborating closely with other departments allows you to create software solutions that are both technically advanced and practically effective. Here, you'll find an environment where creativity and technological excellence go hand in hand. If you're eager to turn ideas into reality and enjoy taking technology to the next level, the Software Development Team at NEURA offers the perfect challenge for you.

Your mission & challenges
  • You are the go-to expert for NEURA's GPU cluster infrastructure – a large‑scale AWS HyperPod environment running cutting‑edge GPU instances for foundation model training and customer fine‑tuning workloads. You design the operational framework, build self‑service tooling for ML teams, and work directly with AWS to influence the platform at the hyperscaler level.

  • Your focus is on cluster engineering and operations — not on ML research itself, but on making sure the people doing that research have rock‑solid, efficient, and accessible infrastructure under them.

  • Setting up, configuring, and continuously evolving NEURA's HyperPod clusters, including HyperPod/Slurm and HyperPod/EKS orchestration models.

  • Designing and implementing strategies for cluster stability: node failure detection, automated job recovery, checkpoint coordination, and fault‑tolerant multi‑node training workflows.

  • Providing a workload priority management framework that allows multiple teams and use cases like foundation model pretraining, fine‑tuning, customer workloads, to share cluster capacity efficiently and fairly.

  • Optimizing end‑to‑end GPU utilization: identifying and resolving bottlenecks across compute, GPU memory, EFA networking, and storage throughput.

  • Working directly and closely with the AWS HyperPod product and solutions engineering teams, escalating operational issues, sharing learnings from one of the platform's largest deployments, and placing concrete requirements on the roadmap.

  • Providing self‑service tooling that allows ML researchers and engineers to launch, monitor, and manage training jobs independently, without requiring infrastructure intervention for routine operations.

  • Developing onboarding documentation, training materials, and internal workshops that enable users to operate efficiently, follow best practices, and understand cost implications of their workloads.

  • Infrastructure as Code is a given for you. Every cluster configuration, every operational change, every new environment is code first.

  • Owning the cost and capacity strategy: Spot instance management, Reserved Instance planning, Savings Plans, and ongoing commitment negotiations with AWS.

What we can look forward to
  • 5+ years of experience in infrastructure or systems engineering, with a strong focus on GPU cluster or HPC operations.

  • Deep hands‑on experience with AWS HyperPod and AWS instances; direct prior experience with HyperPod is a strong differentiator.

  • Solid understanding of both Slurm and Kubernetes as cluster orchestration layers, and the ability to evaluate their trade‑offs for large‑scale GPU workloads.

  • Practical knowledge of distributed training – you understand what affects throughput and how to debug it.

  • Experience building self‑service tooling and operational documentation for technical end users.

  • You make complex infrastructure accessible, not just functional.

  • Strong understanding of cloud cost management at scale: Spot interruption handling, capacity reservations, cost attribution across teams and workloads.

  • Comfort working across organizational boundaries – your primary partners are ML researchers, but you'll also work closely with product, finance, and cloud vendor teams.

  • Strong English communication skills. German is a plus.

What you can look forward to
Creative Freedom and Agility

Enjoy a dynamic, self‑reliant work culture with flat hierarchies, flexible hours, and 30 vacation days. Ideal for those seeking an inspiring professional setting, whether you're starting out or an experienced exec.

Passion for Winning

A passionate and highly skilled team of international experts aiming to redefine robot assistants.

Attractive Compensation

Enjoy a competitive salary package along with exclusive employee discounts.

One Team

Whether it's a summer party or company town hall meetings, we celebrate our successes together.

Professional Growth

Support for your personal and professional development.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Cloud Architect - NeuraGYM (human)
Cloud Architect - NeuraGYM (human)

Atlas Metrics • Gemeindeverwaltungsverband Metzingen

Vor Ort
EUR 70.000 - 100.000
30 vacation days
Exclusive employee discounts
Professional development support
+2
Senior Software Engineer - NEURAGym (human)
Senior Software Engineer - NEURAGym (human)

Atlas Metrics • Gemeindeverwaltungsverband Metzingen

Vor Ort
EUR 60.000 - 90.000
30 vacation days
Employee discounts
Support for personal and professional development
Product Owner (human) - NeuraGym Platform and Integration
Product Owner (human) - NeuraGym Platform and Integration

Atlas Metrics • Gemeindeverwaltungsverband Metzingen

Vor Ort
EUR 65.000 - 85.000
30 vacation days
Competitive salary package
Exclusive employee discounts
+1
Robot Platform Engineer (human)
Robot Platform Engineer (human)

Atlas Metrics • München

Vor Ort
EUR 60.000 - 80.000
Creative work culture
Flexible hours
30 vacation days
+2
Platform Engineer - Neuraverse (human)
Platform Engineer - Neuraverse (human)

Atlas Metrics • München

Vor Ort
EUR 90.000 - 130.000
30 vacation days
Competitive salary
Employee discounts
Forward Deployed Engineer
Forward Deployed Engineer

turbalance • Heidelberg

Hybrid
EUR 60.000 - 80.000
Competitive compensation
Performance-based incentives
Subsidized Deutschlandticket
+2
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Hybrid
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Tech Lead — Robot Systems & Integration (human)
Tech Lead — Robot Systems & Integration (human)

Atlas Metrics • Gemeindeverwaltungsverband Metzingen

Vor Ort
EUR 70.000 - 90.000
30 vacation days
Flexible working hours
Employee discounts
Fullstack Engineer (human)
Fullstack Engineer (human)

Atlas Metrics • Gemeindeverwaltungsverband Metzingen

Vor Ort
EUR 50.000 - 80.000
30 vacation days
Competitive salary package
Exclusive employee discounts
Head of Compute Engineering
Head of Compute Engineering

Impossible Cloud GmbH • Hamburg

Vor Ort
EUR 80.000 - 100.000
Competitive salary
ESOP
Subsidized gym membership
+1