Network CCL Engineer

Evollabs Tech

Dubai

On-site

AED 300,000 - 540,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evollabs Tech in Dubai is seeking an experienced AI infrastructure engineer to design and implement high-performance collective communication for AI accelerators and heterogeneous NPU-GPU systems. You will work at the intersection of accelerator runtime, networking, and distributed AI workloads to enable scalable, low-latency inference across multi-chip rack-scale deployments.

You will develop runtime components for device discovery, rank management, and synchronization, while optimizing

Qualifications

  • Strong C/C++ and Python programming skills with experience building low-level runtime, driver, networking, or distributed systems software.
  • Deep understanding of collective communication primitives (AllReduce, ReduceScatter, AllGather, AllToAll, Broadcast, Send/Recv, Barrier).
  • Experience with accelerator runtime concepts like queues, streams, events, and memory management.

Responsibilities

  • Design and implement collective communication primitives for AI accelerators.
  • Build NPU-native CCL support for multi-chip and rack-scale AI inference systems.
  • Design and implement heterogeneous NPU-GPU communication paths for distributed AI workloads.
  • Develop communication runtime components for device discovery, rank management, topology awareness, and synchronization.
  • Integrate transport-layer across PCIe, Ethernet/RDMA, custom accelerator fabrics.
  • Collaborate with hardware, runtime, compiler, and model-serving teams to co-design scalable communication paths for AI workloads.
  • Profile, debug, and optimize communication bottlenecks including bandwidth and latency.
  • Build tests for multi-device and heterogeneous communication scenarios.

Skills

C/C++
Python
Runtime dev
Distributed systems
Networking
Performance engineering
GPU/ NPU memory

Education

MS/PhD in CS/CE

Tools

NCCL
MPI
UCX
libfabric

Job description

We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference. Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand high-performance networking, collective communication, and distributed AI systems.

This role is focused on designing and implementing collective communication support for our AI accelerator platform, including both NPU-native communication and heterogeneous NPU-GPU communication. You will work at the intersection of accelerator runtime, networking, distributed AI workloads, and hardware-software co-design.

Responsibilities
  • Design and implement collective communication primitives for AI accelerators.
  • Build NPU-native CCL support for multi-chip and rack-scale AI inference systems.
  • Design and implement heterogeneous NPU-GPU communication paths for distributed AI workloads.
  • Develop communication runtime components for device discovery, rank management, topology awareness, stream/queue integration, events, and synchronization.
  • Work on transport-layer integration across PCIe, Ethernet/RDMA, custom accelerator fabrics, or other device-to-device communication mechanisms.
  • Collaborate with hardware, runtime, compiler, and model-serving teams to co-design scalable communication paths for AI workloads.
  • Profile, debug, and optimize communication bottlenecks including bandwidth, latency, congestion, synchronization overhead, and compute-communication overlap.
  • Build correctness, performance, and stress tests for multi-device and heterogeneous communication scenarios.
Requirements
  • 5+ years of relevant experience with M.S./Ph.D. degree in CS/CE or equivalent experience.
  • Strong C/C++ and python programming skills with experience building low-level runtime, driver, networking, or distributed systems software.
  • Deep understanding of collective communication primitives such as AllReduce, ReduceScatter, AllGather, AllToAll, Broadcast, Send/Recv, and Barrier.
  • Familiarity with any networking and distributed communication libraries such as NCCL, RCCL, oneCCL, MPI, UCX, or libfabric.
  • Strong knowledge of GPU/NPU memory models, accelerator memory allocation, memory registration, peer-to-peer transfers, and heterogeneous device communication.
  • Experience with accelerator runtime concepts such as queues, streams, events, command submission, completion handling, synchronization, and memory management.
  • Strong understanding of hardware-software interaction in accelerator systems, including DMA engines, device memory, interrupts, doorbells, fences, and completion signaling.
  • Knowledge of high-performance networking like InfiniBand, and RoCE.
  • Ability to work close to hardware, including DMA engines, queues, streams, fences, completion signaling, NoC/fabric behavior, and accelerator runtime integration.
  • Strong performance engineering skills for optimizing bandwidth, latency, synchronization overhead, congestion, and compute-communication overlap.
Nice to Have
  • Familiarity with PyTorch distributed backends, custom ProcessGroup implementations, or distributed framework integration.
  • Experience optimizing communication for LLM workloads, including tensor parallelism, MoE expert parallelism, KV-cache transfer, or disaggregated inference.
  • Experience integrating communication backends with distributed AI frameworks or serving systems such as PyTorch distributed, vLLM, SGLang, or similar platforms.
What We're Not Looking For

This role is not focused on general backend networking, web services, prompt engineering, or application-layer GenAI development. Instead, it is centered on low-level collective communication, accelerator runtime, heterogeneous device communication, and distributed AI performance at scale.

Why Join Us
  • Work on cutting-edge AI hardware designed specifically for large-scale inference
  • Build core communication infrastructure for next-generation NPU and heterogeneous NPU-GPU systems.
  • Solve challenging problems at the intersection of AI accelerators, networking, distributed runtime, and datacenter-scale systems.
  • Collaborate with hardware, compiler, runtime, and model-serving teams on a full-stack AI infrastructure platform.
  • Make a direct impact on the scalabili
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network CCL Engineer
Network CCL Engineer

Gateworth Group • Dubai

On-site
AED 350,000 - 700,000
Competitive salary
Bonus
Benefits
Senior Network CCL Engineer for AI Accelerators
Senior Network CCL Engineer for AI Accelerators

Gateworth Group • Dubai

On-site
AED 350,000 - 700,000
Competitive salary
Bonus
Benefits
Senior AI Accelerator Networking Engineer
Senior AI Accelerator Networking Engineer

Evollabs Tech • Dubai

On-site
AED 300,000 - 540,000
Runtime Engineer (C/C++)
Runtime Engineer (C/C++)

Evollabs Tech • Dubai

On-site
AED 400,000 - 650,000
RISC-V/NPU Architect
RISC-V/NPU Architect

Evollabs Tech • Dubai

On-site
AED 500,000 - 900,000
Staff Hardware Board Design Engineer
Staff Hardware Board Design Engineer

Evollabs Tech • Dubai

On-site
AED 400,000 - 600,000
High Performance Computing Software Engineer - Supercomputing
High Performance Computing Software Engineer - Supercomputing

Institute of Foundation Models • Abu Dhabi

On-site
AED 350,000 - 700,000
Principal Network Engineer
Principal Network Engineer

Remotedxb • Dubai

On-site
AED 400,000 - 620,000
Senior AI Engineer – LLM Systems
Senior AI Engineer – LLM Systems

Evollabs Tech • Dubai

On-site
AED 420,000 - 660,000
Applied AI/ML Scientist
Applied AI/ML Scientist

Cerebras • United Arab Emirates

On-site
Innovative work culture
Opportunity for continuous learning
Diverse team environment