GPU Cluster Engineer (human)

Atlas Metrics

Riederich

Vor Ort

EUR 90.000 - 120.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Verschicke keinen generischen Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

NEURA Robotics in Metzingen, Germany is seeking a GPU Cluster Engineer to own and optimize NEURA's large-scale GPU infrastructure. You will design, implement, and operate HyperPod clusters, collaborating with ML researchers to ensure robust, scalable training environments.

The role emphasizes cluster engineering, tooling development, and close collaboration with cross-functional teams. German language skills are a plus, and the work is on-site in the Metzingen/Riederich area.

Qualifikationen

  • 5+ years of experience in infrastructure or systems engineering with a focus on GPU cluster or HPC operations.
  • Excellent English communication; German is a plus.
  • Experience with cloud platforms and orchestration tools.
  • Ability to design self-service tooling and documentation for technical users.

Aufgaben

  • You are the go-to expert for NEURA's GPU cluster infrastructure within a large-scale AWS HyperPod environment.
  • Focus on cluster engineering and operations to provide rock-solid infrastructure for ML researchers.
  • Set up, configure, and evolve HyperPod clusters and orchestration models (HyperPod/Slurm, HyperPod/EKS).
  • Develop strategies for cluster stability, automated job recovery, and fault-tolerant multi-node training workflows.
  • Provide a workload priority management framework to share cluster capacity across teams and use cases.

Kenntnisse

Strong English communication

Tools

AWS HyperPod
Slurm
Kubernetes
EKS

Jobbeschreibung

  • NEW
GPU Cluster Engineer (human)

Neura Robotics • Metzingen

Shape the Future of Human-Robot Collaboration

In the Software Department, you're shaping robotic solutions that redefine human-machine collaboration. You'll work with cutting-edge technology, setting industry-changing standards. Not only will you help develop our solutions, but you'll also set new trends and drive innovations forward. In an agile and interdisciplinary team, you'll engage in exciting projects. With clear Scrum processes like daily stand-ups, sprint planning, and reviews, you remain flexible and efficient. Collaborating closely with other departments allows you to create software solutions that are both technically advanced and practically effective. Here, you'll find an environment where creativity and technological excellence go hand in hand. If you're eager to turn ideas into reality and enjoy taking technology to the next level, the Software Development Team at NEURA offers the perfect challenge for you.

Full-time

Metzingen

from today

Your mission & challenges
  • You are the go-to expert for NEURA's GPU cluster infrastructure - a large-scale AWS HyperPod environment running cutting-edge GPU instances for foundation model training and customer fine-tuning workloads. You design the operational framework, build self-service tooling for ML teams, and work directly with AWS to influence the platform at the hyperscaler level.
  • Your focus is on cluster engineering and operations — not on ML research itself, but on making sure the people doing that research have rock-solid, efficient, and accessible infrastructure under them.
  • Setting up, configuring, and continuously evolving NEURA's HyperPod clusters, including HyperPod/Slurm and HyperPod/EKS orchestration models.
  • Designing and implementing strategies for cluster stability: node failure detection, automated job recovery, checkpoint coordination, and fault-tolerant multi-node training workflows.
  • Providing a workload priority management framework that allows multiple teams and use cases like foundation model pretraining, fine-tuning, customer workloads, to share cluster capacity efficiently and fairly.
  • Optimizing end-to-end GPU utilization: identifying and resolving bottlenecks across compute, GPU memory, EFA networking, and storage throughput.
  • Working directly and closely with the AWS HyperPod product and solutions engineering teams, escalating operational issues, sharing learnings from one of the platform's largest deployments, and placing concrete requirements on the roadmap.
  • Providing self-service tooling that allows ML researchers and engineers to launch, monitor, and manage training jobs independently, without requiring infrastructure intervention for routine operations.
  • Developing onboarding documentation, training materials, and internal workshops that enable users to operate efficiently, follow best practices, and understand cost implications of their workloads.
  • Infrastructure as Code is a given for you. Every cluster configuration, every operational change, every new environment is code first.
  • Owning the cost and capacity strategy: Spot instance management, Reserved Instance planning, Savings Plans, and ongoing commitment negotiations with AWS.
What we can look forward to
  • 5+ years of experience in infrastructure or systems engineering, with a strong focus on GPU cluster or HPC operations.
  • Deep hands-on experience with AWS HyperPod and AWS instances; direct prior experience with HyperPod is a strong differentiator.
  • Solid understanding of both Slurm and Kubernetes as cluster orchestration layers, and the ability to evaluate their trade-offs for large-scale GPU workloads.
  • Practical knowledge of distributed training - you understand what affects throughput and how to debug it.
  • Experience building self-service tooling and operational documentation for technical end users.
  • You make complex infrastructure accessible, not just functional.
  • Strong understanding of cloud cost management at scale: Spot interruption handling, capacity reservations, cost attribution across teams and workloads.
  • Comfort working across organizational boundaries — your primary partners are ML researchers, but you'll also work closely with product, finance, and cloud vendor teams.
  • Strong English communication skills. German is a plus.
What you can look forward to
Creative Freedom and Agility

Enjoy a dynamic, self-reliant work culture with flat hierarchies, flexible hours, and 30 vacation days. Ideal for those seeking an inspiring professional setting, whether you're starting out or an experienced exec.

Passion for Winning

A passionate and highly skilled team of international experts aiming to redefine robot assistants.

Attractive Compensation

Enjoy a competitive salary package along with exclusive employee discounts.

One Team

Whether it's a summer party or company town hall meetings, we celebrate our successes together.

Professional Growth

Support for your personal and professional development.

Our values. The cornerstones of our success.
STRONGER TOGETHER

We are a team. We strive to achieve great things by promoting the success of our colleagues and partners.

PASSION DRIVES US

We strive for technological progress in order to give people back their valuable time for enjoyable activities.

MAKING A CHANGE

We strive to revolutionize the world of robotics by pushing the boundaries of technology every day.

TRUST AND HONESTY

We live a high level of appreciation through open communication and transparency.

WE SPEED THINGS UP

We do our best to always be two steps ahead. We achieve this through empowerment, freedom of action and personal responsibility.

WE ARE HUMAN

People are at the center of everything we do.

Our Location

Headquarters: Innovate in Riederich, Live in Metzingen and Stuttgart

Our headquarters in Metzingen and Riederich are the heart of our company. It's not just home to our offices, but also our production facilities, Academy, logistics, and Tech Labs—all working together to turn ideas into reality. Riederich itself is a small, peaceful town, just a kilometer away from Metzingen, a city with its own unique character. Metzingen is globally renowned as Outlet City, attracting visitors from all over the world. Here, you can enjoy exclusive designer stores in a relaxed and charming setting. The city also offers a variety of restaurants, cafés, and a down-to-earth Swabian coziness—perfect for unwinding after work.

Our application process

We ensure a transparent and efficient process and look forward to getting to know you during the application process.

Unsere Mission

David Reger

Gründer und CEO

\"Our goal was to develop the world's first cognitive robot that can work with people, learn from them and provide them with targeted support. And that is exactly what we have achieved. But there is much more to come!\"

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Group Lead System Development (human)
Group Lead System Development (human)

Atlas Metrics • Riederich

Vor Ort
EUR 90.000 - 130.000
30 vacation days
Flexible hours
Flat hierarchies
People Development & Culture Manager (human)
People Development & Culture Manager (human)

Atlas Metrics • Riederich

Vor Ort
EUR 65.000 - 90.000
Flexible hours
30 vacation days
Employee discounts
+1
Production Team Lead - Commissioning of Humanoid Robots (human)
Production Team Lead - Commissioning of Humanoid Robots (human)

Atlas Metrics • Riederich

Vor Ort
EUR 70.000 - 110.000
30 vacation days
Flexible hours
Employee discounts
Robot Operator (human)
Robot Operator (human)

Atlas Metrics • Riederich

Vor Ort
EUR 42.000 - 60.000
30 vacation days
Competitive salary
Cloud Architect - NeuraGYM (human)
Cloud Architect - NeuraGYM (human)

Atlas Metrics • Riederich

Vor Ort
EUR 90.000 - 140.000
Robotics Control Engineer (human)
Robotics Control Engineer (human)

Atlas Metrics • Riederich

Vor Ort
EUR 70.000 - 95.000
30 vacation days
Flexible hours
Employee discounts
Robot Platform Engineer (human)
Robot Platform Engineer (human)

Atlas Metrics • Riederich

Vor Ort
EUR 90.000 - 120.000
Flexible hours
30 vacation days
Robot Perception Expert (human)
Robot Perception Expert (human)

Atlas Metrics • Riederich

Vor Ort
EUR 70.000 - 110.000
30 vacation days
Flexible hours
Exclusive employee discounts
+1
Hardware Architect - Cognitive Robotics (human)
Hardware Architect - Cognitive Robotics (human)

Atlas Metrics • Riederich

Vor Ort
EUR 90.000 - 150.000
30 vacation days
Competitive salary
Flexible hours
Robot Client SDK Engineer (human)
Robot Client SDK Engineer (human)

Atlas Metrics • Riederich

Vor Ort
EUR 65.000 - 90.000