ML Infrastructure Engineer — Multi-Cluster GPU & Training

Prior Labs GmbH

North Canton, New York (OH, NY)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote-friendly
Offsite team events
Global offices Berlin Freiburg NewYork

Job summary

Prior Labs seeks an infrastructure engineer to own and grow its multi-cluster GPU platform. You will manage Slurm on GCP today, plan multi-provider expansion, and optimize scheduling, performance, and costs to empower researchers.

You will collaborate with researchers to understand workloads, reduce bottlenecks, and build the tooling layer and CI pipelines that keep experiments moving rapidly. Remote-optional with travel.

Qualifications

  • 3+ years building and operating production GPU infrastructure at scale.
  • Deep hands-on experience with Slurm and cluster management.
  • Expert-level systems thinking: memory bandwidth and profiling.
  • Strong Python and deep knowledge of PyTorch internals.
  • Track record of improving training throughput or cost efficiency.
  • Fluent AI tooling experience (Claude Code, Cursor, etc.).

Responsibilities

  • Own and evolve multi-cluster GPU infrastructure across providers.
  • Drive GPU utilization and training throughput; identify bottlenecks.
  • Architect next-gen infrastructure: orchestration, new GPUs, capacity planning.
  • Build developer productivity: CI pipelines, experiment tracking, model registry.
  • Own compute budget and optimize cost per FLOP across hardware.

Skills

GPU infrastructure
Python
Systems thinking
PyTorch internals
Performance tuning
AI tooling

Tools

Slurm
GCP
Docker
wandb
GitHub Actions
Triton

Job description

Prior Labs seeks an infrastructure engineer to own and grow its multi-cluster GPU platform. You will manage Slurm on GCP today, plan multi-provider expansion, and optimize scheduling, performance, and costs to empower researchers.

You will collaborate with researchers to understand workloads, reduce bottlenecks, and build the tooling layer and CI pipelines that keep experiments moving rapidly. Remote-optional with travel.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
Staff Engineer, Distributed GPU Clusters & Infra
Staff Engineer, Distributed GPU Clusters & Infra

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Remote HPC Infra Engineer — GPU Clusters
Remote HPC Infra Engineer — GPU Clusters

ElevenLabs • Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Infra Lead: Slurm, Kubernetes & Platform
Senior GPU Infra Lead: Slurm, Kubernetes & Platform

Jobgether SRL • United States

Remote
USD 170,000 - 250,000
ML Infrastructure Engineer – GPU Compute Platform
ML Infrastructure Engineer – GPU Compute Platform

Remanence • Paris (TX)

Hybrid
USD 124,000 - 186,000
Visa sponsorship
Relocation support
Hybrid work setup
+1
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3