A complete application in a minute — tailored resume and cover letter, ready to send.
Evollabs, a technology company designing advanced server hardware for AI workloads, seeks engineers to integrate our AI accelerator backend with modern LLM inference servers and optimize performance across NPU and heterogeneous NPU-GPU systems.
Ideal candidates have 5+ years of experience, strong C/C++ and Python skills, and hands-on work with LLM serving internals. Join us to push the boundaries of AI infrastructure and scalable inference.
We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.
Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand LLM inference servers, model execution runtimes, MoE inference, and distributed AI systems. This role is focused on integrating our AI accelerator platform as a backend into modern LLM inference servers such as vLLM, SGLang, TensorRT-LLM, or similar serving systems.
You will work on adding backend support for our hardware, enabling dense and MoE LLM inference, integrating custom operators and runtime paths, and optimizing the execution of large-scale models across NPU and heterogeneous NPU-GPU systems. You will work at the intersection of LLM serving, accelerator runtime, model execution, distributed inference, and hardware-software co-design.
This role is not for LLM deployment engineers who have only used inference servers to serve models or has written low-level kernels. We are also not looking for candidates whose experience is limited to running, configuring, or deploying models using vLLM, SGLang, TensorRT-LLM, or similar frameworks. We are looking for engineers who have worked on inference server internals. Are familiar with modifying the codebase, adding backend support, integrating custom operators, changing execution paths, or contributing to the framework itself.