Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA

Santa Clara (CA)

On-site

USD 184,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity and benefits

Job summary

NVIDIA is seeking an AI Solutions Architect to design and optimize large-scale GPU-based AI infrastructure. You will collaborate with leading customers to maximize GPU utilization, throughput, and reliability while reducing costs.

The role covers clusters, networking, storage, scheduling, orchestration, and observability, with emphasis on NCCL, InfiniBand, and GPUDirect RDMA. You will work with NVIDIA engineering, product, and sales teams to deliver cutting-edge solutions, run proofs-of-concept,

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or equivalent experience.
  • 6+ years in AI infrastructure, systems engineering, HPC, networking, or site reliability.
  • Deep understanding of Linux systems, distributed computing, GPU architectures and large-scale AI clusters.
  • Hands-on experience designing, deploying and operating high-performance GPU networks on-prem or in the cloud using InfiniBand, RoCE, or GPUDirect RDMA.
  • Experience debugging NCCL communication and distributed performance; topology, transport, routing, and host-level configuration.

Responsibilities

  • Collaborate with customers to maximize GPU utilization and throughput while improving reliability and reducing costs.
  • Design and optimize large-scale AI clusters across compute, networking, storage, scheduling, orchestration and observability.
  • Profile training and inference workloads to identify bottlenecks across GPUs, CPUs, memory and network fabrics.
  • Diagnose complex infra issues spanning InfiniBand, RoCE, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.
  • Lead proofs-of-concept and performance studies, develop benchmarking tools, automation, and runbooks.
  • Partner with engineering, product and sales to secure design wins and drive NVIDIA solutions.

Skills

Python
Linux
Distributed computing
GPU architectures
Networking
Kubernetes
Slurm
Shell scripting

Education

BS/MS/PhD in CS/EE/Physics/Math

Tools

NCCL
InfiniBand/RoCE
NCCL tests
DCGM
Nsight Systems
NVLink/NVSwitch

Job description

NVIDIA is seeking an AI Solutions Architect to design and optimize large-scale GPU-based AI infrastructure. You will collaborate with leading customers to maximize GPU utilization, throughput, and reliability while reducing costs.

The role covers clusters, networking, storage, scheduling, orchestration, and observability, with emphasis on NCCL, InfiniBand, and GPUDirect RDMA. You will work with NVIDIA engineering, product, and sales teams to deliver cutting-edge solutions, run proofs-of-concept,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Infra Architect for Enterprise GPU Clusters
Lead AI Infra Architect for Enterprise GPU Clusters

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Senior AI Infrastructure Architect — Enterprise GPU Clusters
Senior AI Infrastructure Architect — Enterprise GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Architect – GPU Clusters
Senior AI Infrastructure Architect – GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infra Architect for Enterprise GPU Clusters
Senior AI Infra Architect for Enterprise GPU Clusters

NVIDIA • New York (NY)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Solutions Architect — GPU Cloud
Senior AI Solutions Architect — GPU Cloud

NVIDIA • Virginia (MN)

On-site
USD 184,000 - 288,000
Senior AI Solutions Architect — GPU Cloud Infra
Senior AI Solutions Architect — GPU Cloud Infra

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Health benefits
Senior Networking Solutions Architect – AI Clusters
Senior Networking Solutions Architect – AI Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
AI Infrastructure Solutions Architect - GPU Data Center
AI Infrastructure Solutions Architect - GPU Data Center

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity options
Benefits package
Cloud Solutions Architect — AI & GPU Infrastructure
Cloud Solutions Architect — AI & GPU Infrastructure

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Factory Architect: GPUs, NCCL, Automation, Equity
Senior AI Factory Architect: GPUs, NCCL, Automation, Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits