Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA is seeking a GPU System/Fabrics Architect in Bengaluru to design and validate multi-GPU scale-up/down systems for AI datacenters. You will couple GPU compute with high-bandwidth memory, in-package interconnects, and GPU-to-GPU fabrics to optimize performance, resilience, and scalability.
The role involves architecture definition for NVLink, Ethernet and other interconnects, with emphasis on HW/SW co-design and cross-functional collaboration across ASIC, compiler, and software teams.
NVIDIA has continuously reinvented itself. Our invention of the GPU sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionised parallel computing. Today, research in artificial intelligence is booming worldwide, which calls for highly scalable and massively parallel computation horsepower that NVIDIA GPUs excel. NVIDIA has continuously reinvented itself. Our invention of the GPU sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionised parallel computing. Today, research in artificial intelligence is booming worldwide, which calls for highly scalable and massively parallel computation horsepower that NVIDIA GPUs excel. We are seeking a GPU System/Fabrics Architect who will architect and design multi-GPU scale-up and scale-out systems for next-generation AI datacenter platforms. The architect in this role will explore and define architectures that tightly couple GPU compute, high-bandwidth memory, in-package interconnects, and GPU-to-GPU communication Fabric transport/routing subsystems to deliver industry-leading AI performance, scalability, and resilience.
Architect multi-GPU systems for scale-up and scale-out configurations, balancing AI performance, scalability, and resilience for the Agentic era. Define, modify, and evaluate future architectures for high-speed interconnects such as NVLink and Ethernet co-designed with the GPU memory system and networking hardware. Architect RDMA-capable hardware and define transport layer optimizations for GPU-based large scale AI workload deployments. Explore and build novel high-density multi-chiplet, multi-package, multi-node rack-scale AI systems consisting of hundreds/thousands of copper and optically interconnected GPUs. Use and modify system models, perform simulations, and bottleneck analyses to guide design trade-offs. Work with GPU ASIC, compiler, library, and software teams to enable efficient hardware-software co-design across compute, memory, and communication layers.
BS/MS/PhD in Electrical Engineering, Computer Engineering, or equivalent area. 2+ years or more of relevant experience in system design and/or ASIC/SoC architecture for GPU, CPU, XPU, or networking products. Deep understanding of communication interconnect protocols such as Ethernet, InfiniBand, NVLink, CXL and PCIe. Proven ability to architect multi-GPU/multi-CPU topologies, with awareness of bandwidth scaling, NUMA, memory models, coherency, and resilience. Strong analytical and system modeling skills for performance, power, resilience. Excellent cross-functional collaboration and skills.
Experience with NICs, DPUs, RDMA/RoCE or InfiniBand transport offload architectures. Expertise in chiplet interconnect architectures or multi-node fabrics and protocols for high-performance distributed computing.
#LI-Hybrid NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.