We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference. Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand high-performance networking, collective communication, and distributed AI systems.
This role is focused on designing and implementing collective communication support for our AI accelerator platform, including both NPU-native communication and heterogeneous NPU-GPU communication. You will work at the intersection of accelerator runtime, networking, distributed AI workloads, and hardware-software co-design.
Responsibilities
- Design and implement collective communication primitives for AI accelerators.
- Build NPU-native CCL support for multi-chip and rack-scale AI inference systems.
- Design and implement heterogeneous NPU-GPU communication paths for distributed AI workloads.
- Develop communication runtime components for device discovery, rank management, topology awareness, stream/queue integration, events, and synchronization.
- Work on transport-layer integration across PCIe, Ethernet/RDMA, custom accelerator fabrics, or other device-to-device communication mechanisms.
- Collaborate with hardware, runtime, compiler, and model-serving teams to co-design scalable communication paths for AI workloads.
- Profile, debug, and optimize communication bottlenecks including bandwidth, latency, congestion, synchronization overhead, and compute-communication overlap.
- Build correctness, performance, and stress tests for multi-device and heterogeneous communication scenarios.
Requirements
- 5+ years of relevant experience with M.S./Ph.D. degree in CS/CE or equivalent experience.
- Strong C/C++ and python programming skills with experience building low-level runtime, driver, networking, or distributed systems software.
- Deep understanding of collective communication primitives such as AllReduce, ReduceScatter, AllGather, AllToAll, Broadcast, Send/Recv, and Barrier.
- Familiarity with any networking and distributed communication libraries such as NCCL, RCCL, oneCCL, MPI, UCX, or libfabric.
- Strong knowledge of GPU/NPU memory models, accelerator memory allocation, memory registration, peer-to-peer transfers, and heterogeneous device communication.
- Experience with accelerator runtime concepts such as queues, streams, events, command submission, completion handling, synchronization, and memory management.
- Strong understanding of hardware-software interaction in accelerator systems, including DMA engines, device memory, interrupts, doorbells, fences, and completion signaling.
- Knowledge of high-performance networking like InfiniBand, and RoCE.
- Ability to work close to hardware, including DMA engines, queues, streams, fences, completion signaling, NoC/fabric behavior, and accelerator runtime integration.
- Strong performance engineering skills for optimizing bandwidth, latency, synchronization overhead, congestion, and compute-communication overlap.
Nice to Have
- Familiarity with PyTorch distributed backends, custom ProcessGroup implementations, or distributed framework integration.
- Experience optimizing communication for LLM workloads, including tensor parallelism, MoE expert parallelism, KV-cache transfer, or disaggregated inference.
- Experience integrating communication backends with distributed AI frameworks or serving systems such as PyTorch distributed, vLLM, SGLang, or similar platforms.
What We're Not Looking For
This role is not focused on general backend networking, web services, prompt engineering, or application-layer GenAI development. Instead, it is centered on low-level collective communication, accelerator runtime, heterogeneous device communication, and distributed AI performance at scale.
Why Join Us
- Work on cutting-edge AI hardware designed specifically for large-scale inference
- Build core communication infrastructure for next-generation NPU and heterogeneous NPU-GPU systems.
- Solve challenging problems at the intersection of AI accelerators, networking, distributed runtime, and datacenter-scale systems.
- Collaborate with hardware, compiler, runtime, and model-serving teams on a full-stack AI infrastructure platform.
- Make a direct impact on the scalabili