Principal Engineer – Scale-Up GPU Networking (HPC / AI)This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.**Who We Are:**Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.**Job Description:****High Performance Computing, AI and Labs** is a critical element of HPE. We are focused on delivering innovative solutions that accelerate our customers’ digital transformation, enabling them to tackle their complex, and data-intensive workloads. Combining deep expertise and the development of the world’s most cutting-edge, high-performance supercomputers, is defining the next era of computing delivering valuable insight & innovation. Join us and redefine what’s next for you. # **What you'll do:**## ## **Key Responsibilities** ### **Architect & Deliver Scale-Up Networking*** Design and implement **GPU-aware networking paths** for high-bandwidth, low-latency intra-node communication.* Develop and optimize **GPU → NIC → GPU** data movement, shared memory models, and DMA pathways.### **GPU Ecosystem Integration*** Work with **NVIDIA CUDA, NVLink, NCCL**, and **AMD ROCm, InfinityFabric, RCCL** teams to integrate and optimize scale-up communication semantics.* Drive improvements to **DMA engines, BAR mappings, ATS/IOMMU**, and GPU memory registration workflows.### **Runtime & Communication Stack Development*** Enhance and extend **Libfabric, UCX, CXI, SHMEMX, OpenMPI** for GPU-accelerated scale-up workflows.* Optimize communication collectives, transport layers, and GPU-direct capabilities.### **Multi-NIC / NUMA Performance Optimization*** Characterize and tune **multi-NIC per socket**, NUMA-zone mapping, GPU locality, CQ/queue design, and CPU/GPU topology optimization.### **Upstreaming & Architecture Influence*** Lead upstream contributions to open-source projects (OFI, UCX, OpenMPI, RCCL/NCCL enablement).* Partner with HPC/AI ecosystem teams to shape future architectures.### **Debugging, Performance, and Quality*** Own complex debugging across **driver, runtime, GPU, kernel, and user-space** boundaries.* Develop profiling workflows using Nsight, ROCm tools, eBPF, perf, etc.# # **What you need to bring:**## ## **Required Skills & Experience** * 10–15+ years building **high-performance networking, GPU, or kernel-level software**.* Deep expertise in **C/C++**, Linux internals, memory management, RDMA, PCIe, IOMMU, ATS, DMA engines.* Strong understanding of **CUDA, ROCm, GPU memory models, P2P, GDS (GPUDirect Storage), GDR (GPUDirect RDMA)**.* Hands-on experience with **MPI, SHMEM, Libfabric, UCX**, or similar communication stacks.* Proven experience driving **architecture**, cross-org technical decisions, and upstream contributions.* Ability to mentor senior engineers, influence multi-team designs, and own end-to-end delivery.## ## **Preferred Qualifications*** Experience with **NIC architecture** (CXI, RoCE, Infiniband, Slingshot, NVLink Switch).* Experience optimizing **collectives (AllReduce/AllGather)** on GPUs.* Background contributing to **open-source HPC/AI libraries**.* Familiarity with **HPC system architecture**, NUMA tuning, and multi-accelerator systems.