An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Evollabs Tech is developing advanced AI accelerator hardware and software for large-scale inference. The team seeks engineers who understand LLM inference servers, runtime paths, and distributed AI systems to integrate our backend into modern serving frameworks.
You will optimize inference execution, support device registration, and extend operator paths across NPU-GPU deployments, contributing to performance and reliability at datacenter scale.
We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.
Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand LLM inference servers, model execution runtimes, MoE inference, and distributed AI systems. This role is focused on integrating our AI accelerator platform as a backend into modern LLM inference servers such as vLLM, SGLang, TensorRT-LLM, or similar serving systems.
You will work on adding backend support for our hardware, enabling dense and MoE LLM inference, integrating custom operators and runtime paths, and optimizing the execution of large-scale models across NPU and heterogeneous NPU-GPU systems. You will work at the intersection of LLM serving, accelerator runtime, model execution, distributed inference, and hardware-software co-design.
This role is not for LLM deployment engineers who have only used inference servers to serve models or has written low-level kernels. We are also not looking for candidates whose experience is limited to running, configuring, or deploying models using vLLM, SGLang, TensorRT-LLM, or similar frameworks. We are looking for engineers who have worked on inference server internals. Are familiar with modifying the codebase, adding backend support, integrating custom operators, changing execution paths, or contributing to the framework itsel