Principal Engineer – Scale-Up GPU Networking (HPC / AI)

Hewlett Packard Enterprise Development LP

Bengaluru

Hybrid

INR 6,500,000 - 9,000,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health & Wellbeing
Personal & Professional Development
Unconditional Inclusion

Job summary

Hewlett Packard Enterprise Development LP in Bengaluru is seeking a Principal Engineer to lead scale-up GPU networking for HPC/AI workloads. Design, implement, and optimize high-bandwidth, low-latency intra-node communication paths using GPU tech and open-source stacks.

The role requires deep expertise in C/C++, Linux internals, RDMA, and CUDA/ROCm, with a track record of end-to-end delivery and upstream contributions. Hybrid work model applies.

Qualifications

  • 10-15+ years building high-performance networking, GPU, or kernel-level software.
  • Deep expertise in C/C++, Linux internals, memory management, RDMA, PCIe, IOMMU, DMA engines.
  • Strong understanding of CUDA, ROCm, GPU memory models, P2P, GPUDirect.
  • Hands-on experience with MPI, SHMEM, Libfabric, UCX, or similar stacks.
  • Proven experience driving architecture, cross-org decisions, and upstream contributions.
  • Ability to mentor senior engineers, influence multi-team designs, and own end-to-end delivery.

Responsibilities

  • Architect and deliver scale-up GPU networking design for high bandwidth, low-latency intra-node paths.
  • Develop and optimize GPU→NIC→GPU data movement, shared memory models, and DMA pathways.
  • Integrate with CUDA, NVLink, NCCL, ROCm and InfinityFabric; improve DMA and memory registration workflows.
  • Extend Libfabric/UCX/CXI/SHMEMX/OpenMPI for GPU-accelerated scale-up workflows.
  • Tune multi-NIC per socket, NUMA zones, GPU locality, and topology for performance.
  • Lead upstream contributions to OFI/UCX/OpenMPI; collaborate with HPC/AI teams on future architectures.
  • Own debugging across driver, runtime, GPU, kernel, and user-space; develop profiling workflows.

Skills

HPC Networking
GPU kernel software
C/C++
Linux internals
RDMA
MPI/UCX

Tools

MPI
SHMEM
Libfabric
UCX
NCCL
CUDA
ROCm
NVLink

Job description

Principal Engineer – Scale-Up GPU Networking (HPC / AI)

Hybrid: Work on average 2 days per week from an HPE office.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description:

High Performance Computing, AI and Labs is a critical element of HPE. We are focused on delivering innovative solutions that accelerate our customers’ digital transformation, enabling them to tackle their complex, and data-intensive workloads. Combining deep expertise and the development of the world’s most cutting-edge, high-performance supercomputers, is defining the next era of computing delivering valuable insight & innovation. Join us and redefine what’s next for you.

What you'll do:
  • Key Responsibilities:
    • Architect & Deliver Scale-Up Networking Design and implement GPU‑aware networking paths for high‑bandwidth, low‑latency intra‑node communication.
    • Develop and optimize GPU → NIC → GPU data movement, shared memory models, and DMA pathways.
  • GPU Ecosystem Integration: Work with NVIDIA CUDA, NVLink, NCCL, and AMD ROCm, InfinityFabric, RCCL teams to integrate and optimize scale‑up communication semantics. Drive improvements to DMA engines, BAR mappings, ATS/IOMMU, and GPU memory registration workflows.
  • Runtime & Communication Stack Development: Enhance and extend Libfabric, UCX, CXI, SHMEMX, OpenMPI for GPU‑accelerated scale‑up workflows. Optimize communication collectives, transport layers, and GPU‑direct capabilities.
  • Multi‑NIC / NUMA Performance Optimization: Characterize and tune multi‑NIC per socket, NUMA‑zone mapping, GPU locality, CQ/queue design, and CPU/GPU topology optimization.
  • Upstreaming & Architecture Influence: Lead upstream contributions to open‑source projects (OFI, UCX, OpenMPI, RCCL/NCCL enablement). Partner with HPC/AI ecosystem teams to shape future architectures.
  • Debugging, Performance, and Quality: Own complex debugging across driver, runtime, GPU, kernel, and user‑space boundaries. Develop profiling workflows using Nsight, ROCm tools, eBPF, perf, etc.
What you need to bring:
Required Skills & Experience
  • 10-15+ years building high‑performance networking, GPU, or kernel‑level software.
  • Deep expertise in C/C++, Linux internals, memory management, RDMA, PCIe, IOMMU, ATS, DMA engines.
  • Strong understanding of CUDA, ROCm, GPU memory models, P2P, GDS (GPUDirect Storage), GDR (GPUDirect RDMA).
  • Hands‑on experience with MPI, SHMEM, Libfabric, UCX, or similar communication stacks.
  • Proven experience driving architecture, cross‑org technical decisions, and upstream contributions.
  • Ability to mentor senior engineers, influence multi‑team designs, and own end‑to‑end delivery.
Preferred Qualifications
  • Experience with NIC architecture (CXI, RoCE, Infiniband, Slingshot, NVLink Switch).
  • Experience optimizing collectives (AllReduce/AllGather) on GPUs.
  • Background contributing to open‑source HPC/AI libraries.
  • Familiarity with HPC system architecture, NUMA tuning, and multi‑accelerator systems.
Accessibility

HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here. Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability.

What We Can Offer You:
  • Health & Wellbeing: We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.
  • Personal & Professional Development: We also invest in your career because the better you are, the better we all are.
  • Unconditional Inclusion: We are unconditionally inclusive in the way we work and celebrate individual uniqueness.
Job:

Engineering Job Level: TCP_05

Equal Employment Opportunity

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together.

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

HPE is an E-Verify employer.

Equal Opportunity Employer (EEO) Hewlett Packard Enterprise provides equal employment opportunity to any employee or applicant without regard to sex, gender, color, race, ethnicity, religion, creed, national origin, ancestry, citizenship, age, marital status, sexual orientation, gender identity and expression, physical or mental disability, medical condition, pregnancy, protected veteran status, uniformed service status, familial status, genetic information, political affiliation, or any other characteristic protected under federal, state, or local law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Engineer - GPU Networking for High Performance Computing and AI
Principal Engineer - GPU Networking for High Performance Computing and AI

Hewlett Packard Enterprise India Private Limited • Bengaluru

Hybrid
INR 4,000,000 - 6,000,000
Hybrid work model
Inclusive culture
Professional development
Principal Engineer – Scale-Up GPU Networking (HPC / AI)
Principal Engineer – Scale-Up GPU Networking (HPC / AI)

Hewlett Packard Enterprise Development LP • India

Hybrid
INR 5,000,000 - 8,500,000
Principal Software Engineer- SDET
Principal Software Engineer- SDET

Hewlett Packard Enterprise • Bengaluru

On-site
INR 1,800,000 - 2,600,000
Sr Network Test Engineer — Data Center Fabrics
Sr Network Test Engineer — Data Center Fabrics

Hewlett Packard Enterprise Development LP • India

On-site
INR 1,800,000 - 3,000,000
Performance Engineering Architect
Performance Engineering Architect

Hewlett Packard Enterprise • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Health & Wellbeing
Career development programs
Unconditional inclusion
Sr Network Test Engineer — Data Center Fabrics
Sr Network Test Engineer — Data Center Fabrics

Hewlett Packard Enterprise Development LP • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Health benefits
Career development
Inclusive culture
Sr Hardware Engineer _ PCB/System Design
Sr Hardware Engineer _ PCB/System Design

Hewlett Packard Enterprise • Bengaluru

Hybrid
INR 1,500,000 - 2,100,000
Hardware Engineer III - PCB Design
Hardware Engineer III - PCB Design

Hewlett Packard Enterprise • Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Performance Engineering Architect
Performance Engineering Architect

Hewlett Packard Enterprise India Private Limited • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Hardware Engineer - PCB and System Design
Senior Hardware Engineer - PCB and System Design

Hewlett Packard Enterprise India Private Limited • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Health & Wellbeing
Personal & Professional Development
Unconditional Inclusion