ML engineer (LLM quantization & optimization)

ODS Serbia

Time

On-site

NOK 900,000 - 1,300,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Sobolev Research Center is seeking an ML Engineer to develop and optimize algorithms that accelerate large language model inference, directly impacting latency, cost efficiency, and scalability of production-grade AI systems.

You will explore and implement cutting-edge techniques such as speculative decoding, prompt compression, quantization, and generation optimisation to improve performance across deployment scenarios.

Qualifications

  • Strong experience with deep learning frameworks (PyTorch or TensorFlow).
  • Solid understanding of Transformer architectures and LLMs.
  • Experience with model inference optimization.
  • Strong Python skills.
  • Understanding of GPU/CPU performance and memory bottlenecks.

Responsibilities

  • LLM Quantization & Low-Precision Optimization: develop and apply quantization techniques to reduce model memory footprint and accelerate inference while preserving output quality.
  • Speculative Decoding & Generation Acceleration: design algorithms to reduce decoding steps and improve generation speed.
  • Prompt Compression & Context Optimization: reduce input context length without degrading output quality.

Skills

Transformer basics
Python programming
Model inference optimization
GPU/CPU performance

Tools

PyTorch
TensorFlow
ONNX Runtime
Hugging Face Transformers
TensorRT
DeepSpeed
BitsAndBytes
GPTQ
FlashAttention
xFormers

Job description

We are looking for an ML Engineer to focus on developing and optimizing algorithms that accelerate large language model (LLM) inference. Your work will directly impact latency, cost efficiency, and scalability of production-grade AI systems. You will explore and implement cutting-edge techniques such as speculative decoding, prompt compression, quantization, and generation optimisation.


About The Company

Company Sobolev Research Center. Sobolev Research Center is a local branch of an international full cycle IT company, based in Novosibirsk.


Responsibilities


  • LLM Quantization & Low-Precision Optimization (AWQ, GPTQ, SmoothQuant, BitsAndBytes, INT8/INT4/FP8) Develop and apply quantization techniques to reduce model memory footprint and accelerate inference while preserving output quality. What you will do: Implement and evaluate weight-only and weight-activation quantization methods, including AWQ, GPTQ, SmoothQuant, and related approaches. Analyze how different quantization algorithms work and compare their accuracy, latency, memory usage, calibration requirements, and hardware efficiency. Work with symmetric and asymmetric quantization, as well as per-tensor, per-channel, and group-wise quantization schemes. Understand and apply both post-training quantization (PTQ) and quantization-aware training (QAT), selecting the appropriate approach for different models and deployment scenarios.

  • Speculative Decoding & Generation Acceleration (EAGLE3, dFlash, dSpark, dFly, MTP) Design algorithms that reduce the number of decoding steps and improve generation speed. What you will do: a. Implement speculative decoding pipelines (draft + target models). b. Develop multi-token prediction approaches. c. Explore parallel and tree-based decoding strategies.

  • Prompt Compression & Context Optimization (Token pruning / attention-based filtering, semantic compression via embeddings, LLM-based summarization (self-compression)) Reduce input context length without degrading output quality. What you will do: a. Compress long prompts and conversation history. b. Filter irrelevant tokens dynamically. c. Optimize context window usage.


Requirements


  • Strong experience with deep learning frameworks (PyTorch or TensorFlow).

  • Solid understanding of Transformer architectures and LLMs.

  • Experience with model inference optimization.

  • Strong Python skills.

  • Understanding of GPU/CPU performance and memory bottlenecks. [Tech Stack] PyTorch, Hugging Face Transformers, TensorRT, ONNX Runtime, vLLM, SGLang, DeepSpeed, FlashAttention, xFormers, Quantization tools (BitsAndBytes, GPTQ).


Working Conditions

Full time, office mode only, Novosibirsk, VMI.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

In-Office LLM Inference Engineer — Quantization
In-Office LLM Inference Engineer — Quantization

ODS Serbia • Time

On-site
NOK 900,000 - 1,300,000
Senior AI Engineer - LLM, Multimodal, High-Load (Remote)
Senior AI Engineer - LLM, Multimodal, High-Load (Remote)

ODS Serbia • Time

On-site
NOK 199,000 - 353,000
Медицинское страхование
Корпоративная связь
Льготы по cafeteria-системе
+1
Research Engineer (LLM Performance), London
Research Engineer (LLM Performance), London

Isomorphic Labs • London

Hybrid
NOK 877,000 - 1,504,000
Engineering Manager | Inference
Engineering Manager | Inference

United States Digital Space LLC • London

Hybrid
NOK 1,331,000 - 1,996,000
Hybrid work
Hack Fridays
30 days annual leave
+1
Senior AI Engineer
Senior AI Engineer

ODS Serbia • Time

On-site
NOK 199,000 - 353,000
Медицинское страхование
Корпоративная связь
Льготы по cafeteria-системе
+1
Python Developer (AI/LLM)
Python Developer (AI/LLM)

Halliburton • Oslo

On-site
NOK 60,000 - 80,000
Competitive salaries and pension schemes
Outstanding insurance coverage
Discounts on recreational activities
AI/Harness Engineer (Python)
AI/Harness Engineer (Python)

Tenth Revolution Group • Oslo

On-site
NOK 900,000 - 1,200,000
Member of Technical Staff, Applied AI
Member of Technical Staff, Applied AI

Latent Labs • London

Hybrid
NOK 1,557,000 - 2,466,000
Private health insurance
Pension contributions
Generous leave policies (including par
+1
AI Data Science Intern (UK)
AI Data Science Intern (UK)

TWG Global AI • London

On-site
NOK 589,000 - 710,000
Forward-Deployed AI Engineer for Biotech Deployments
Forward-Deployed AI Engineer for Biotech Deployments

Latent Labs • London

Hybrid
NOK 973,000 - 1,428,000
Private health insurance
Pension contributions
Generous leave policies
+2