Principal engineer – GPU Memory Systems & Scale-Up/Out Systems

METAVERSE COMPUTING LIMITED

Hong Kong

On-site

HKD 1,200,000 - 1,600,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

METAVERSE COMPUTING LIMITED is seeking a Principal engineer to push the boundaries of GPU memory systems and scale-up/out architectures. You will design memory hierarchies, interconnects, and distributed systems for AI/ML and HPC workloads, validating ideas with simulators or hardware prototypes.

The role requires a PhD or equivalent, deep expertise in computer architecture and memory systems, and a track record of publications in prestigious venues.

Qualifications

  • PhD in CS/EE or related field.
  • (or equivalent research experience).
  • Strong background in computer architecture, memory systems, and parallel/distributed computing.
  • Hands-on experience with GPU architectures, CUDA, RDMA, or high-performance interconnects (CXL).
  • Proficiency in performance modeling and simulation.
  • Strong programming skills in C/C++, Python, or SystemVerilog/HDL.
  • Track record of publications in ISCA, MICRO, ASPLOS, HPCA, SC, or related venues.

Responsibilities

  • Research and develop novel GPU memory architectures (cache hierarchies, near-memory computing, disaggregated memory, link/switch optimizations).
  • Design and evaluate scale-up and scale-out systems for GPU clusters, focusing on interconnect topologies, communication protocols, and load balancing.
  • Optimize memory bandwidth, latency, and capacity utilization for large-scale GPU workloads.
  • Collaborate with hardware and software teams to prototype new architectures (simulators, FPGA emulation, or real hardware).
  • Publish cutting-edge research in top conferences/journals and contribute to patents.
  • Engage with industry and academic partners to drive innovation in GPU-accelerated computing.

Skills

GPU architectures
CUDA
RDMA
CXL
Performance modeling
Parallel computing
Distributed computing
SystemVerilog/HDL
Python
C/C++

Education

PhD in Computer Science, Electrical Engineering, or related field

Tools

GPGPU-Sim
Gem5
NS3

Job description

Principal engineer – GPU Memory Systems & Scale-Up/Out Systems

We are seeking a highly motivated **Principal engineer** to advance the state-of-the-art in **GPU memory systems** and **scale-up/out computing architectures**. This role involves designing and optimizing memory hierarchies, interconnects, and distributed systems to improve performance, efficiency, and scalability for next-generation GPU-accelerated workloads (e.g., AI/ML, HPC, and large-scale data analytics).


The ideal candidate will have deep expertise in **computer architecture, memory systems, parallel computing, and distributed systems**, with a strong publication record in top-tier conferences (e.g., ISCA, MICRO, ASPLOS, HPCA, SC, NeurIPS).


Key Responsibilities


  • Research and develop novel **GPU memory architectures** (e.g., cache hierarchies, near-memory computing, disaggregated memory, link/switch optimizations).

  • Design and evaluate **scale-up and scale-out systems** for GPU clusters, focusing on **interconnect topologies, communication protocols, and load balancing**.

  • Optimize **memory bandwidth, latency, and capacity utilization** for large-scale GPU workloads.

  • Collaborate with hardware and software teams to prototype new architectures (e.g., using simulators, FPGA emulation, or real hardware).

  • Publish cutting-edge research in top conferences/journals and contribute to patents.

  • Engage with industry and academic partners to drive innovation in GPU-accelerated computing.


Required Qualifications


  • **PhD in Computer Science, Electrical Engineering, or related field** (or equivalent research experience).

  • Strong background in **computer architecture, memory systems, and parallel/distributed computing**.

  • Hands-on experience with **GPU architectures, CUDA, RDMA, or high-performance interconnects (CXL)**.

  • Proficiency in **performance modeling and simulation** (e.g., GPGPU-Sim, SST, Gem5, NS3).

  • Strong programming skills in **C/C++, Python, or SystemVerilog/HDL** for prototyping.

  • Track record of publications in **ISCA, MICRO, ASPLOS, HPCA, SC, or related venues**.


Preferred Qualifications


  • Experience with **disaggregated memory, cache coherence protocols, or memory-centric computing**.

  • Knowledge of **AI/ML workloads and their memory/system bottlenecks**.

  • Contributions to **open-source hardware/software projects** (e.g., LLVM, PyTorch, OpenMPI).

  • Familiarity with **datacenter-scale systems** (e.g., distributed training, high-performance storage).


Why Join Us?


  • Work on **cutting-edge research** with real-world impact in AI, HPC, and cloud computing.

  • Collaborate with **world-class researchers and engineers**.

  • Competitive compensation, equity, and publication/patent incentives.


This job description balances technical depth with broad applicability, attracting candidates with expertise in **GPU memory systems** and **large-scale computing**. Let me know if you'd like any refinements!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal GPU Memory Systems Engineer — Scale-Up/Out
Principal GPU Memory Systems Engineer — Scale-Up/Out

METAVERSE COMPUTING LIMITED • Hong Kong

On-site
HKD 1,200,000 - 1,600,000
GPU Server Hardware Validation Engineer
GPU Server Hardware Validation Engineer

Pragmatike • Hong Kong

On-site
HKD 420,000 - 700,000
Server Engineer (HK - Hybrid)
Server Engineer (HK - Hybrid)

Pragmatike • Hong Kong

Hybrid
HKD 480,000 - 720,000
Systems Engineer
Systems Engineer

Tower Research Capital • Hong Kong

On-site
HKD 450,000 - 750,000
Generous PTO
Savings plans
Hybrid work
+6
Head of AI Infrastructure (Supercomputing Center)
Head of AI Infrastructure (Supercomputing Center)

Captiare Limited • Hong Kong Island

On-site
HKD 1,200,000 - 1,800,000
Inference & System Optimization Engineer (Experienced )
Inference & System Optimization Engineer (Experienced )

奇瑞全球創新(香港)有限公司 • Hong Kong Island

On-site
HKD 900,000 - 1,300,000
Server Engineer (On-site Hong Kong)
Server Engineer (On-site Hong Kong)

Pragmatike • Hong Kong

On-site
HKD 420,000 - 660,000
Senior AI Infrastructure Architect — GPU Clusters & Security
Senior AI Infrastructure Architect — GPU Clusters & Security

CPJobs International • Hong Kong

On-site
HKD 600,000 - 1,000,000
Platform Engineer - Tribus
Platform Engineer - Tribus

Tribus • Hong Kong

On-site
HKD 480,000 - 900,000
Principal Core Infrastructure Engineer – GPU
Principal Core Infrastructure Engineer – GPU

Ll Oefentherapie • Hong Kong

On-site
HKD 900,000 - 1,300,000