GPU Cluster Architect

Nebius B.V.

India

On-site

INR 4,000,000 - 6,500,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nebius B.V. is seeking a GPU Cluster Architect to drive end-to-end architectural decisions for a next-generation AI infrastructure in India.

You will design scalable GPU cluster topologies, investigate interconnects, storage, and control planes, and ensure performance and reliability across multiple data centers. The role is highly hands-on, requiring collaboration with site reliability, networking, storage, and DC engineering teams, with a focus on training workloads, scalability, and

Qualifications

  • 5+ years of experience designing clusters.
  • Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.).
  • Experience with HPC interconnects (InfiniBand & RoCE).
  • Solid background in systems architecture, networking, and hardware reliability.
  • Experience in scripting for automation and telemetry pipelines (Python, Go).

Responsibilities

  • Cluster Design: Architect scalable GPU cluster topologies including compute nodes, interconnect, storage, and control planes.
  • Performance Modeling: Analyze AI/ML workloads to inform design tradeoffs across latency, bandwidth, and GPU density.
  • Network Architecture: Validate low-latency, high-throughput interconnects at POD and DC scale.
  • Storage Integration: Optimize performance for training datasets and checkpointing.
  • Reliability & Monitoring: Analyze signals from monitoring systems to detect design issues.
  • Collaboration: Partner with site reliability, networking, storage, and DC engineering teams to operationalize and scale the architecture.

Skills

GPU architecture
Cluster design
HPC interconnects
Networking
Automation scripting

Tools

Python
Go

Job description

We are seeking a GPU Cluster Architect to drive the design of our next-generation AI infrastructure. In this high-impact, hands-on role, you will make end-to-end architectural decisions across compute, networking, and storage — ensuring our platforms can meet the massive scale, performance, and reliability requirements of modern AI workloads.

This is a high-impact, hands-on architecture role where you’ll define how tens of thousands of GPUs are interconnected, cooled down, powered, and optimized across multiple data center sites.

Responsibilities
  • Cluster Design: Architect scalable GPU cluster topologies including compute nodes, interconnect (InfiniBand, Ethernet), storage, and control planes.
  • Performance Modeling: Analyze AI/ML workloads (e.g. LLM training, inference) to inform design tradeoffs across latency, bandwidth, and GPU density.
  • Network Architecture: Align with network architect relevant design and validate low-latency, high-throughput interconnects (e.g., InfiniBand HDR/NDR, RoCEv2) at POD and DC scale.
  • Storage Integration: Work with storage teams to optimize performance for training datasets, checkpointing, and others.
  • Reliability & Monitoring: Understand and analyze signal from monitoring systems to the detect flows in design.
  • Collaboration: Partner with site reliability, networking, storage, and DC engineering teams to operationalize and scale your architecture.
Requirements
  • 5+ years of experience designing clusters.
  • Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.).
  • Experience with HPC interconnects (InfiniBand & RoCE).
  • Solid background in systems architecture, networking, and hardware reliability.
  • Experience in scripting for automation and telemetry pipelines (Python, Go, etc.).
  • 5+ years of experience designing clusters, Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.), Experience with HPC interconnects (InfiniBand & RoCE), Solid background in systems architecture, networking, and hardware reliability, Experience in scripting for automation and telemetry pipelines (Python, Go, etc.)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Architect
GPU Architect

NVIDIA Corporation • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Corporation • India

On-site
INR 3,500,000 - 7,000,000
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
Principal Software Architect- High Performance Computing
Principal Software Architect- High Performance Computing

Applied Materials India • Chennai District

On-site
INR 4,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Corporation • Bengaluru

On-site
INR 4,000,000 - 7,000,000
GPU Cluster Architect
GPU Cluster Architect

Nebius • Shigevadi

On-site
INR 4,000,000 - 7,000,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
AI Compute Engineer
AI Compute Engineer

DAMAC Digital • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Lead Engineer (HPC, GPU, CUDA)
Lead Engineer (HPC, GPU, CUDA)

AIRA Matrix • Thane

On-site
INR 1,200,000 - 2,500,000