ML Engineer, Infrastructure

Meyandy LLC

Berlin

Hybrid

EUR 70.000 - 110.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Meyandy LLC is seeking an experienced Infrastructure Engineer to own and evolve multi-cluster GPU infrastructure, collaborating with research teams to accelerate discovery.

You’ll drive GPU utilization, profile bottlenecks, optimize memory and communications, and plan capacity across providers for scalable performance.

Qualifikationen

  • 3+ years building and operating production GPU infrastructure at scale.
  • Experience with Slurm and cluster management.
  • Strong scripting and debugging skills for distributed training.

Aufgaben

  • Own and evolve multi-cluster GPU infrastructure across providers.
  • Drive GPU utilization and training throughput.
  • Architect the next generation of infrastructure: multi-cluster orchestration, new GPU generations, provider diversification, capacity planning.
  • Build the developer productivity layer: CI pipelines, experiment tracking, model registry, data processing, and internal tooling.
  • Own the compute budget and optimize cost per FLOP across providers and hardware.

Kenntnisse

GPU infrastructure
Distributed training systems
Cost optimization

Tools

Slurm
GCP
Docker
wandb
GitHub Actions
PyTorch
Triton

Jobbeschreibung

## Who we areFoundation models transformed text and images. Structured data - the largest and most consequential data format in the world - stayed untouched, until now. What LLMs did for language, we're doing for tables.We pioneered tabular foundation models: TabPFN v2 was a Nature cover story, has passed 3.5M+ downloads and 7,500+ GitHub stars, and runs in production from detecting lung disease with Oxford Cancer Analytics to preventing train failures with Hitachi. The hardest problems - millions of rows, real-time inference, entirely new modalities - are still open, and no one else is working on them at this level.We're a small, highly selective team of 40+ with backgrounds from Google, DeepMind, Meta, Apple, Amazon, Jane Street, and CERN, led by Frank Hutter, Noah Hollmann, and Sauraj Gambhir, and advised by Bernhard Schölkopf and Turing Award winner Yann LeCun.In July 2026, less than 18 months after our €9M pre-seed, we joined SAP as an independent frontier AI lab - same team, mission, and open-weights models, now backed by more than €1 billion over four years.## **About the Role**We spend tens of millions per year on GPU compute to train tabular foundation models. That's not a target, it's what we're running today, and it's growing. The person who owns this infrastructure makes decisions worth millions of dollars: cluster architecture, scheduling efficiency, provider strategy, hardware selection. A wrong call costs six figures.Today we run Slurm on GCP across multiple clusters. We're scaling to multi-cluster, multi-provider infrastructure and evaluating new hardware generations as they come online. You own the full stack, from cluster operations and cost optimization to distributed training performance and the tooling layer that keeps researchers moving fast. You work directly with the research team and understand what they're doing well enough to make infrastructure decisions that actually help them. And this isn't a pure support role. We operate an open environment. If you've got the next SOTA tabular architecture up your sleeve, go ahead and train it.**What you'll work on:*** Own and evolve multi-cluster GPU infrastructure. Slurm on GCP today, multi-provider and new hardware tomorrow. Architecture, scheduling, reliability, cost optimization* Drive GPU utilization and training throughput: profiling, memory optimization, communication bottlenecks, systems-level debugging of distributed training across large runs* Architect the next generation of our infrastructure: multi-cluster orchestration, new GPU generations, provider diversification, capacity planning against growing compute demands* Build the developer productivity layer: CI pipelines, experiment tracking, model registry, data processing, and internal tooling that keeps research iteration speed high* Own the compute budget. You understand cost per FLOP across providers and hardware, and you hate wasted compute**Tech stack:** Slurm, GCP, Docker, wandb, GitHub Actions, uv, PyTorch, Triton**You may be a good fit if you have:*** 3+ years building and operating production GPU infrastructure or distributed training systems at scale. At a major AI lab, a well-funded ML startup, or an HPC environment* Deep hands-on experience with Slurm and cluster man...
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

ML Engineer, Infrastructure
ML Engineer, Infrastructure

DUDE CHEM • Berlin

Vor Ort
EUR 90.000 - 135.000
Full Stack Engineer, ML Platform
Full Stack Engineer, ML Platform

Meyandy LLC • Berlin

Hybrid
EUR 70.000 - 110.000
Research Scientist, Foundational Data Science
Research Scientist, Foundational Data Science

Meyandy LLC • Berlin

Hybrid
EUR 80.000 - 120.000
Research Scientist Intern (PhD)
Research Scientist Intern (PhD)

Meyandy LLC • Berlin

Hybrid
EUR 70.000 - 120.000
Mentorship
Compute resources
Healthcare
+2
GPU Cluster Engineer - Scalable AI Infrastructure
GPU Cluster Engineer - Scalable AI Infrastructure

NEURA Robotics • Deutschland

Remote
USD 140.000 - 195.000
Machine Learning Systems & Infrastructure Engineer
Machine Learning Systems & Infrastructure Engineer

SpAItial • München

Vor Ort
EUR 70.000 - 90.000
Senior Forward Deployed Platform Engineer
Senior Forward Deployed Platform Engineer

Recare • Berlin

Hybrid
EUR 90.000 - 130.000
Flexible work location
Office in Berlin Mitte
Learning budget
+2
GPU Cluster Engineer (human)
GPU Cluster Engineer (human)

NEURA Robotics • Deutschland

Vor Ort
USD 140.000 - 195.000
Research Scientist, Foundational Data Science
Research Scientist, Foundational Data Science

Prior Labs • Berlin

Vor Ort
EUR 70.000 - 110.000
ML Engineer – MLOps & Platform Engineering (m/w/d)
ML Engineer – MLOps & Platform Engineering (m/w/d)

Remotely • Würselen

Vor Ort
EUR 90.000 - 130.000
Remote work from Germany
Occasional travel for team events