GPU Cluster Architect

Nebius B.V.

India

On-site

INR 4,000,000 - 6,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebius B.V. is seeking a GPU Cluster Architect to drive end-to-end architectural decisions for a next-generation AI infrastructure in India.

You will design scalable GPU cluster topologies, investigate interconnects, storage, and control planes, and ensure performance and reliability across multiple data centers. The role is highly hands-on, requiring collaboration with site reliability, networking, storage, and DC engineering teams, with a focus on training workloads, scalability, and

Qualifications

  • 5+ years of experience designing clusters.
  • Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.).
  • Experience with HPC interconnects (InfiniBand & RoCE).
  • Solid background in systems architecture, networking, and hardware reliability.
  • Experience in scripting for automation and telemetry pipelines (Python, Go).

Responsibilities

  • Cluster Design: Architect scalable GPU cluster topologies including compute nodes, interconnect, storage, and control planes.
  • Performance Modeling: Analyze AI/ML workloads to inform design tradeoffs across latency, bandwidth, and GPU density.
  • Network Architecture: Validate low-latency, high-throughput interconnects at POD and DC scale.
  • Storage Integration: Optimize performance for training datasets and checkpointing.
  • Reliability & Monitoring: Analyze signals from monitoring systems to detect design issues.
  • Collaboration: Partner with site reliability, networking, storage, and DC engineering teams to operationalize and scale the architecture.

Skills

GPU architecture
Cluster design
HPC interconnects
Networking
Automation scripting

Tools

Python
Go

Job description

We are seeking a GPU Cluster Architect to drive the design of our next-generation AI infrastructure. In this high-impact, hands-on role, you will make end-to-end architectural decisions across compute, networking, and storage — ensuring our platforms can meet the massive scale, performance, and reliability requirements of modern AI workloads.

This is a high-impact, hands-on architecture role where you’ll define how tens of thousands of GPUs are interconnected, cooled down, powered, and optimized across multiple data center sites.

Responsibilities
  • Cluster Design: Architect scalable GPU cluster topologies including compute nodes, interconnect (InfiniBand, Ethernet), storage, and control planes.
  • Performance Modeling: Analyze AI/ML workloads (e.g. LLM training, inference) to inform design tradeoffs across latency, bandwidth, and GPU density.
  • Network Architecture: Align with network architect relevant design and validate low-latency, high-throughput interconnects (e.g., InfiniBand HDR/NDR, RoCEv2) at POD and DC scale.
  • Storage Integration: Work with storage teams to optimize performance for training datasets, checkpointing, and others.
  • Reliability & Monitoring: Understand and analyze signal from monitoring systems to the detect flows in design.
  • Collaboration: Partner with site reliability, networking, storage, and DC engineering teams to operationalize and scale your architecture.
Requirements
  • 5+ years of experience designing clusters.
  • Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.).
  • Experience with HPC interconnects (InfiniBand & RoCE).
  • Solid background in systems architecture, networking, and hardware reliability.
  • Experience in scripting for automation and telemetry pipelines (Python, Go, etc.).
  • 5+ years of experience designing clusters, Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.), Experience with HPC interconnects (InfiniBand & RoCE), Solid background in systems architecture, networking, and hardware reliability, Experience in scripting for automation and telemetry pipelines (Python, Go, etc.)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU System Architect
Senior GPU System Architect

NVIDIA Gruppe • Bengaluru

On-site
INR 9,578,000 - 14,368,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Corporation • India

On-site
INR 3,500,000 - 7,000,000
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
GPU Cluster Architect
GPU Cluster Architect

Nebius • Shigevadi

On-site
INR 4,000,000 - 7,000,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Gruppe • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Gruppe • Bengaluru

On-site
INR 400,000 - 900,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000