Inference Server Engineer

Evollabs

Hyderabad

On-site

INR 4,000,000 - 6,500,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Evollabs, a technology company designing advanced server hardware for AI workloads, seeks engineers to integrate our AI accelerator backend with modern LLM inference servers and optimize performance across NPU and heterogeneous NPU-GPU systems.

Ideal candidates have 5+ years of experience, strong C/C++ and Python skills, and hands-on work with LLM serving internals. Join us to push the boundaries of AI infrastructure and scalable inference.

Qualifications

  • 5+ years of relevant experience with MS/Ph.D. degree in CS/CE or equivalent experience.
  • Strong C/C++ and Python programming skills.
  • Hands-on experience modifying, extending, or contributing to any known LLM inference servers such as vLLM, SGLang, TensorRT-LLM etc.
  • Experience running and scaling workloads on large-scale, heterogeneous clusters (CPU + GPU) using distributed training or inference strategies.

Responsibilities

  • Integrate our AI accelerator backend into LLM inference servers such as vLLM, SGLang, TensorRT-LLM, or similar frameworks.
  • Implement backend support for device registration, runtime execution, memory management, custom operators, and KV cache handling.
  • Enable dense and MoE LLM inference on our hardware, including attention execution, expert routing, expert execution, and distributed inference support.
  • Profile, debug, and optimize inference performance while building correctness, performance, and stress tests for NPU and heterogeneous NPU-GPU deployments.

Skills

C/C++
Python
LLM inference servers
Distributed inference
PyTorch internals

Education

MS/PhD in CS/CE

Tools

vLLM
SGLang
TensorRT-LLM

Job description

We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.

Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand LLM inference servers, model execution runtimes, MoE inference, and distributed AI systems. This role is focused on integrating our AI accelerator platform as a backend into modern LLM inference servers such as vLLM, SGLang, TensorRT-LLM, or similar serving systems.

You will work on adding backend support for our hardware, enabling dense and MoE LLM inference, integrating custom operators and runtime paths, and optimizing the execution of large-scale models across NPU and heterogeneous NPU-GPU systems. You will work at the intersection of LLM serving, accelerator runtime, model execution, distributed inference, and hardware-software co-design.

Responsibilities
  • Integrate our AI accelerator backend into LLM inference servers such as vLLM, SGLang, TensorRT-LLM, or similar frameworks.
  • Implement backend support for device registration, runtime execution, memory management, custom operators, and KV cache handling.
  • Enable dense and MoE LLM inference on our hardware, including attention execution, expert routing, expert execution, and distributed inference support.
  • Profile, debug, and optimize inference performance while building correctness, performance, and stress tests for NPU and heterogeneous NPU-GPU deployments.
Requirements
  • 5+ years of relevant experience with M.S./Ph.D. degree in CS/CE or equivalent experience.
  • Strong C/C++ and Python programming skills.
  • Hands-on experience modifying, extending, or contributing to any known LLM inference servers such as vLLM, SGLang, TensorRT-LLM etc.
  • Experience running and scaling workloads on large-scale, heterogeneous clusters (CPU + GPU) using distributed training or inference strategies.
  • Familiarity with PyTorch internals, custom op registration, tensor execution, model loading, and Hugging Face model integration.
  • Familiarity with distributed inference techniques such as tensor parallelism, expert parallelism, data parallelism, multi-device execution, and multi-node serving.
Nice to Have
  • Prior merged PRs or significant contributions to vLLM, SGLang, TensorRT-LLM or similar inference server repositories.
  • Strong knowledge of accelerator runtime and compiler stack.
  • Experience with GPU/NPU memory models, allocation, movement, and heterogeneous device execution.
  • Familiarity with distributed communication libraries such as NCCL, RCCL, oneCCL, MPI, UCX, or libfabric.
  • Experience with CUDA / Triton Kernels Experience
What We're Not Looking For

This role is not for LLM deployment engineers who have only used inference servers to serve models or has written low-level kernels. We are also not looking for candidates whose experience is limited to running, configuring, or deploying models using vLLM, SGLang, TensorRT-LLM, or similar frameworks. We are looking for engineers who have worked on inference server internals. Are familiar with modifying the codebase, adding backend support, integrating custom operators, changing execution paths, or contributing to the framework itself.

Why Join Us?
  • Work on cutting-edge AI hardware designed specifically for large-scale inference.
  • Build core inference server backend infrastructure for next-generation NPU and heterogeneous NPU-GPU systems.
  • Solve challenging problems at the intersection of LLM serving, AI accelerators, distributed runtime, MoE inference, and datacenter-
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

eBay • Bengaluru

On-site
INR 4,000,000 - 7,500,000
MTS 2, AI Platform Professional
MTS 2, AI Platform Professional

The Networker • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Performance Engineer - Inference
Performance Engineer - Inference

Keka Inc. • Bengaluru

On-site
INR 300,000 - 540,000
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru

Vibehackers • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Travel up to 30%
Open-source contributions
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

On-site
INR 4,200,000 - 6,300,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Chennai District

On-site
INR 4,000,000 - 7,000,000
Hybrid work model
Inference Systems Engineer
Inference Systems Engineer

Nava • Bengaluru

On-site
INR 1,700,000 - 2,500,000
Machine Learning Engineer
Machine Learning Engineer

Valiance Solutions • Bengaluru Urban

On-site
INR 3,500,000 - 5,500,000
AI20P Library Engineer, Machine Learning Acceleration
AI20P Library Engineer, Machine Learning Acceleration

Qpisemi • Bengaluru

On-site
INR 4,000,000 - 7,000,000