Technical Product Manager - Soperator

Nebius

Amsterdam

On-site

EUR 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Career growth and learning opportunities
Flexibility and ownership
Collaborative and innovative culture

Job summary

Nebius in Amsterdam is seeking a Technical Product Manager to lead product direction for Soperator, their Slurm-on-Kubernetes control plane for GPU clusters. In this role, you will drive customer discovery, define product metrics, and ensure high-impact results.

You will need 3-5 years of experience in product management within ML infrastructure and strong familiarity with distributed systems. Enjoy a competitive salary and opportunities for career growth in a dynamic AI environment.

Qualifications

  • 3-5+ years in Product Management, ML infrastructure/MLOps, distributed systems, or cloud platform engineering.
  • Strong technical depth in distributed systems, cloud infrastructure, or ML platforms.
  • Hands-on familiarity with large-scale ML training and orchestration tools.

Responsibilities

  • Own the user journey across Soperator clusters: Slurm workflows, dashboards, alerts, and notifications.
  • Define product direction from problem discovery to delivery.
  • Lead deep customer discovery to identify high-impact opportunities.

Skills

Product Management
Distributed systems
ML infrastructure/MLOps
Cloud platform engineering
Stakeholder management

Tools

Slurm
Kubernetes
Ray

Job description

The role

At Nebius, we’re building a next-generation AI compute platform for large-scale ML training and inference — from a few nodes to thousands of GPUs. We’re looking for a Technical Product Manager to own product direction for Soperator — our Slurm-on-Kubernetes control plane for GPU clusters. In this role, you will shape how ML engineers and research teams run, scale, and optimize distributed workloads in production. If you care about systems that combine performance, reliability, and developer experience at the frontier of AI infrastructure, this role is for you.

Your responsibilities will include
  • Own the full user journey across Soperator clusters: Slurm workflows, dashboards, alerts/notifications, node lifecycle, and training/inference capacity management.
  • Define product direction end-to-end: problem discovery to solution design to delivery to adoption.
  • Lead deep customer discovery through interviews, usage analytics, and workload analysis to uncover high-impact opportunities.
  • Drive execution across platform teams: compute, networking, storage, observability, IAM and others.
  • Translate frontier ML and infrastructure ideas into practical product capabilities for real-world GPU clusters.
  • Define success metrics, prioritize roadmap decisions with data, and ensure measurable customer/business impact.
  • Lead the open-source strategy and execution for Soperator: shape public roadmap themes, prioritize OSS-facing capabilities, and ensure strong adoption in the community.
We expect you to have
  • 3-5+ years in Product Management, ML infrastructure/MLOps, distributed systems, or cloud platform engineering.
  • Strong technical depth in distributed systems, cloud infrastructure, or ML platforms.
  • Hands‑on familiarity with large-scale ML training and orchestration tools (e.g. Slurm, Kubernetes, Ray).
  • Track record of shipping technically complex products with multiple engineering teams.
  • Strong communication and stakeholder management across engineering, research, and customers.
  • Experience with product analytics, data-informed prioritization, and experimentation.
  • High ownership, high learning velocity, and comfort operating in fast-moving AI infrastructure environments.
It will be an added bonus if you have
  • Experience with GPU platforms and HPC primitives: InfiniBand/RDMA, topology‑aware scheduling, high-throughput storage.
  • Practical understanding of modern ML training stacks: PyTorch, DeepSpeed, FSDP/ZeRO, NCCL.
  • Familiarity with efficiency and reliability metrics: Goodput, MFU, failure modes, preemption handling, health checks.
  • Exposure to large-scale LLM training/inference systems.
  • Experience in observability, performance tuning, or SRE/reliability engineering.
  • Customer-facing technical experience (solutioning, support, architecture advisory).
Benefits & Perks
  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams
What's it like to work at Nebius
  • Fast moving
  • Bold thinking
  • Constant growth
  • Meaningful impact
  • Trust and real ownership
  • Opportunity to shape the future of AI
Equal Opportunity Statement

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Product Manager - Soperator
Technical Product Manager - Soperator

United States Digital Space LLC • Amsterdam

Hybrid
EUR 80,000 - 110,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Technical Product Manager — AI Infra & GPU Clusters
Technical Product Manager — AI Infra & GPU Clusters

Nebius • Amsterdam

On-site
EUR 70,000 - 90,000
Competitive compensation
Career growth and learning opportunities
Flexibility and ownership
+1
Technical Product Manager – AI Compute Platform
Technical Product Manager – AI Compute Platform

ApplyMint • Netherlands

Hybrid
EUR 70,000 - 100,000
Competitive compensation
Career growth opportunities
Flexible work environment
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

Nebius • Amsterdam

On-site
EUR 70,000 - 90,000
Competitive compensation
Career growth opportunities
Collaborative culture
+1
Staff Backend Engineer / Tech Lead Manager
Staff Backend Engineer / Tech Lead Manager

Neura Market • Amsterdam

On-site
EUR 120,000 - 170,000
Competitive pay
Career growth
Flexible work
+3
Senior Technical Product Manager, Token Factory
Senior Technical Product Manager, Token Factory

Embedded Shishya • Netherlands

Hybrid
EUR 90,000 - 130,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Head of Platform
Head of Platform

AI Chopping Block • Amsterdam

Hybrid
EUR 140,000 - 190,000
Competitive compensation
Career growth and learning
Ownership and autonomy
+1
Head of Platform
Head of Platform

Nebius • Amsterdam

On-site
EUR 150,000 - 210,000
Competitive compensation
Career growth
Flexibility & ownership
+3
Senior Technical Product Manager, Token Factory
Senior Technical Product Manager, Token Factory

Nebius • Amsterdam

On-site
EUR 90,000 - 150,000
Competitive compensation
Career growth opportunities
Flexibility and ownership
+2
Technical Project Manager (Hardware)
Technical Project Manager (Hardware)

Nebius • Amsterdam

On-site
EUR 90,000 - 130,000
Competitive pay
Career growth
Flexibility
+3