Principal engineer – GPU Memory Systems & Scale-Up/Out Systems
We are seeking a highly motivated **Principal engineer** to advance the state-of-the-art in **GPU memory systems** and **scale-up/out computing architectures**. This role involves designing and optimizing memory hierarchies, interconnects, and distributed systems to improve performance, efficiency, and scalability for next-generation GPU-accelerated workloads (e.g., AI/ML, HPC, and large-scale data analytics).
The ideal candidate will have deep expertise in **computer architecture, memory systems, parallel computing, and distributed systems**, with a strong publication record in top-tier conferences (e.g., ISCA, MICRO, ASPLOS, HPCA, SC, NeurIPS).
Key Responsibilities
- Research and develop novel **GPU memory architectures** (e.g., cache hierarchies, near-memory computing, disaggregated memory, link/switch optimizations).
- Design and evaluate **scale-up and scale-out systems** for GPU clusters, focusing on **interconnect topologies, communication protocols, and load balancing**.
- Optimize **memory bandwidth, latency, and capacity utilization** for large-scale GPU workloads.
- Collaborate with hardware and software teams to prototype new architectures (e.g., using simulators, FPGA emulation, or real hardware).
- Publish cutting-edge research in top conferences/journals and contribute to patents.
- Engage with industry and academic partners to drive innovation in GPU-accelerated computing.
Required Qualifications
- **PhD in Computer Science, Electrical Engineering, or related field** (or equivalent research experience).
- Strong background in **computer architecture, memory systems, and parallel/distributed computing**.
- Hands-on experience with **GPU architectures, CUDA, RDMA, or high-performance interconnects (CXL)**.
- Proficiency in **performance modeling and simulation** (e.g., GPGPU-Sim, SST, Gem5, NS3).
- Strong programming skills in **C/C++, Python, or SystemVerilog/HDL** for prototyping.
- Track record of publications in **ISCA, MICRO, ASPLOS, HPCA, SC, or related venues**.
Preferred Qualifications
- Experience with **disaggregated memory, cache coherence protocols, or memory-centric computing**.
- Knowledge of **AI/ML workloads and their memory/system bottlenecks**.
- Contributions to **open-source hardware/software projects** (e.g., LLVM, PyTorch, OpenMPI).
- Familiarity with **datacenter-scale systems** (e.g., distributed training, high-performance storage).
Why Join Us?
- Work on **cutting-edge research** with real-world impact in AI, HPC, and cloud computing.
- Collaborate with **world-class researchers and engineers**.
- Competitive compensation, equity, and publication/patent incentives.
This job description balances technical depth with broad applicability, attracting candidates with expertise in **GPU memory systems** and **large-scale computing**. Let me know if you'd like any refinements!