Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.
Nebius seeks a Technical Product Manager to own the product direction for Soperator — our Slurm-on-Kubernetes control plane for GPU clusters. You will shape how ML engineers and research teams run, scale, and optimize distributed workloads in production.
You will own the full user journey across Soperator clusters, define end-to-end roadmaps, and lead customer discovery, analytics, and data‑informed prioritization.
We are seeking a Technical Product Manager to own the product direction for Soperator — our Slurm-on-Kubernetes control plane for GPU clusters. In this role, you will shape how ML engineers and research teams run, scale, and optimize distributed workloads in production.
Your responsibilities include owning the full user journey across Soperator clusters (Slurm workflows, dashboards, alerts/notifications, node lifecycle, training/inference capacity management), defining product direction end-to-end (problem discovery to delivery and adoption), leading customer discovery through interviews and analytics, coordinating with platform teams (compute, networking, storage, observability, IAM, etc.), translating frontier ML and infrastructure ideas into practical product capabilities, defining success metrics and roadmaps, and leading the OSS strategy for Soperator to drive community adoption.
Requirements include 3–5+ years in Product Management, ML infrastructure/MLOps, distributed systems, or cloud platform engineering; strong technical depth in distributed systems or cloud infrastructure; hands‑on familiarity with large-scale ML training/orchestration tools (Slurm, Kubernetes, Ray); track record shipping technically complex products; strong communication and stakeholder management; experience with product analytics and data‑informed prioritization; high ownership and fast learning velocity.
Bonus points for GPU platforms/HPC primitives experience, modern ML training stacks (PyTorch, DeepSpeed, FSDP/ZeRO, NCCL), observability and SRE/reliability engineering, and customer-facing technical experience.
Benefits & Perks include
About Nebius Nebius AI provides an AI cloud platform with large GPU capacities in Europe; we operate data centers and have a global footprint with R&D hubs across Europe, the UK, North America and Israel.
What it’s like to work at Nebius emphasizes fast‑moving, bold thinking and meaningful impact, with a culture of trust and ownership.
We are an equal opportunity employer and welcome applicants from all backgrounds.