We are looking for an ML Engineer to focus on developing and optimizing algorithms that accelerate large language model (LLM) inference. Your work will directly impact latency, cost efficiency, and scalability of production-grade AI systems. You will explore and implement cutting-edge techniques such as speculative decoding, prompt compression, quantization, and generation optimisation.
About The Company
Company Sobolev Research Center. Sobolev Research Center is a local branch of an international full cycle IT company, based in Novosibirsk.
Responsibilities
- LLM Quantization & Low-Precision Optimization (AWQ, GPTQ, SmoothQuant, BitsAndBytes, INT8/INT4/FP8) Develop and apply quantization techniques to reduce model memory footprint and accelerate inference while preserving output quality. What you will do: Implement and evaluate weight-only and weight-activation quantization methods, including AWQ, GPTQ, SmoothQuant, and related approaches. Analyze how different quantization algorithms work and compare their accuracy, latency, memory usage, calibration requirements, and hardware efficiency. Work with symmetric and asymmetric quantization, as well as per-tensor, per-channel, and group-wise quantization schemes. Understand and apply both post-training quantization (PTQ) and quantization-aware training (QAT), selecting the appropriate approach for different models and deployment scenarios.
- Speculative Decoding & Generation Acceleration (EAGLE3, dFlash, dSpark, dFly, MTP) Design algorithms that reduce the number of decoding steps and improve generation speed. What you will do: a. Implement speculative decoding pipelines (draft + target models). b. Develop multi-token prediction approaches. c. Explore parallel and tree-based decoding strategies.
- Prompt Compression & Context Optimization (Token pruning / attention-based filtering, semantic compression via embeddings, LLM-based summarization (self-compression)) Reduce input context length without degrading output quality. What you will do: a. Compress long prompts and conversation history. b. Filter irrelevant tokens dynamically. c. Optimize context window usage.
Requirements
- Strong experience with deep learning frameworks (PyTorch or TensorFlow).
- Solid understanding of Transformer architectures and LLMs.
- Experience with model inference optimization.
- Strong Python skills.
- Understanding of GPU/CPU performance and memory bottlenecks. [Tech Stack] PyTorch, Hugging Face Transformers, TensorRT, ONNX Runtime, vLLM, SGLang, DeepSpeed, FlashAttention, xFormers, Quantization tools (BitsAndBytes, GPTQ).
Working Conditions
Full time, office mode only, Novosibirsk, VMI.