Remote GPU Cluster Architect for Scalable AI Infrastructure

Nebius

United States

On-site

USD 184,000 - 318,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance: 100% company-paid coverage
401(k) plan: Up to 4% company match
20 weeks paid parental leave for primary caregivers
Remote work reimbursement up to $85/month
Life insurance coverage
Competitive compensation
Career growth opportunities
Collaborative culture

Job summary

Nebius is seeking a GPU Cluster Architect to drive the design of next-generation AI infrastructure. This role involves making critical architectural decisions across compute, networking, and storage, ensuring the infrastructure meets the needs of modern AI workloads.

You will define how GPUs are interconnected and optimized across data center sites. This position supports remote work from the USA and offers competitive salaries ranging from $184K to $318K, including bonuses.

Qualifications

  • 5+ years of experience designing clusters.
  • Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.).
  • Experience with HPC interconnects (InfiniBand & RoCE).
  • Solid background in systems architecture, networking, and hardware reliability.
  • Experience in scripting for automation and telemetry pipelines (Python, Go, etc.).

Responsibilities

  • Architect scalable GPU cluster topologies including compute nodes, interconnect, storage, and control planes.
  • Analyze AI/ML workloads to inform design tradeoffs across latency, bandwidth, and GPU density.
  • Align with network architect design and validate low-latency, high-throughput interconnects at scale.
  • Work with storage teams to optimize performance for training datasets.
  • Understand and analyze signals from monitoring systems to detect flaws in design.
  • Partner with various teams to operationalize and scale architecture.

Skills

Cluster Design
Performance Modeling
Network Architecture
Storage Integration
Reliability & Monitoring
Collaboration
Deep understanding of modern GPU architecture
Experience with HPC interconnects
Scripting for automation

Job description

Nebius is seeking a GPU Cluster Architect to drive the design of next-generation AI infrastructure. This role involves making critical architectural decisions across compute, networking, and storage, ensuring the infrastructure meets the needs of modern AI workloads.

You will define how GPUs are interconnected and optimized across data center sites. This position supports remote work from the USA and offers competitive salaries ranging from $184K to $318K, including bonuses.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Center GPU Infrastructure Engineer
Senior Data Center GPU Infrastructure Engineer

Nebius • New Jersey

On-site
USD 85,000 - 140,000
Career growth opportunities
Flexible work environment
Collaborative culture
+1
Senior AI Cloud Network Engineer
Senior AI Cloud Network Engineer

Nebius • United States

On-site
USD 125,000 - 180,000
Comprehensive health insurance
401(k) plan with company contribution
Paid time off and public holidays
+2
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
GPU Infra Solutions Architect for Large-Scale AI Clusters
GPU Infra Solutions Architect for Large-Scale AI Clusters

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior AI Infra Architect - GPU Clusters & NVLink (Remote)
Senior AI Infra Architect - GPU Clusters & NVLink (Remote)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior GenAI Solutions Architect - Remote & GPU Infrastructure
Senior GenAI Solutions Architect - Remote & GPU Infrastructure

NVIDIA • California (MO)

Hybrid
USD 184,000 - 357,000
Senior Data Center Engineer - GPU/AI Infra
Senior Data Center Engineer - GPU/AI Infra

Nebius B.V. • Vineland (NJ)

On-site
USD 85,000 - 140,000
Career growth
Flexibility and ownership
Collaborative culture
+2
Lead HPC Cluster Engineer for GPU AI Compute
Lead HPC Cluster Engineer for GPU AI Compute

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Lead AI Infrastructure Solutions Architect for HPC Clusters
Lead AI Infrastructure Solutions Architect for HPC Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000