Senior Applied Scientist, Efficient LLM Inference & Model Optimization

Nebius Group

Amsterdam

On-site

EUR 120,000 - 180,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Career growth
Flexible work and ownership
Collaborative culture
Impactful AI projects
International environment

Job summary

Nebius, Amsterdam-based AI cloud company, is seeking scientists to advance frontier LLM and VLM inference bottlenecks into research questions, experiments, and production capabilities. You will own focused research and optimization projects, write strong code, and produce prototypes, publications, and artifacts for engineers to build on.

Your work spans quantization, distillation, speculative decoding, KV-cache optimization, and model/runtime co-optimization, evaluating trade-offs in quality,

Qualifications

  • PhD in computer science, machine learning, ML systems, or a closely related discipline.
  • Strong publication record or equivalent research artifacts in ML/inference/production-serving.
  • Strong Python and PyTorch implementation skills to turn ideas into experiments and prototypes.
  • Deep knowledge of LLMs, VLMs, transformer inference, decoding algorithms, and production-serving trade-offs.
  • Strong experimental design skills covering ablations, baselines, metrics, statistics, and failure analysis.
  • Excellent written and verbal communication.

Responsibilities

  • Lead research projects in efficient LLM and VLM inference, from hypotheses and experiments through ablations, prototypes, and production handoff.
  • Develop and evaluate methods spanning quantization, quantization-aware training (QAT), distillation, speculative decoding, KV-cache reuse, and model/runtime co-optimization.
  • Build prototypes using PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks and collaborate with engineers to turn them into production components.
  • Research LLM request routing and scheduling strategies, including cache-aware load balancing and prefill-decode disaggregation (PDD).
  • Investigate multi-node inference for dense and mixture-of-experts models, including wide expert parallelism and expert placement.
  • Develop rigorous evaluations covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.
  • Share results through internal reports, technical blogs, papers, and open-source artifacts, and mentor engineers and scientists.

Skills

Python & PyTorch
LLMs / VLMs knowledge
Research experiments
Publication record

Education

PhD in CS/ML/related field

Tools

PyTorch
CUDA
Triton

Job description

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role

Nebius Token Factoryis looking forscientists who can turn frontierLLM and VLMinference bottlenecks into researchquestions, conductrigorous experiments, andtranslate their findings into production capabilities. You will own focusedresearch and optimizationprojects, write strong code, and produce prototypes, publications, and technical artifactsthat engineers canbuild on.

Your research will span quantization, distillation, speculative decoding, KV-cache optimization, and model/runtime co-optimization, alongside inference engines and distributed inference architectures. You will investigate how these approaches interact, evaluate their trade-offs across model quality, latency, throughput, memory footprint, and cost per token, and work with engineering teams to bring promising results into production.

Your responsibilities
  • Leadresearch projects in efficient LLM and VLMinference, from hypotheses and experiments through ablations, prototypes, and production handoff.
  • Develop and evaluate methods spanning quantization, quantization-aware training (QAT),distillation, speculative decoding, KV-cache reuse, KV-cache compression, long-context inference, MoE routing, and model/runtime co-optimization.
  • Build prototypesusingPyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks,and collaboratewith MLEs and platform engineers toturn them intoproduction components.
  • Research LLM request routing and scheduling strategies, including cache-aware load balancing and prefill-decode disaggregation (PDD). Evaluate how queueing, KV-cache transfer, worker placement, and prefill/decode capacity allocation affect latency, throughput, and serving cost.
  • Investigate multi-node inference for dense and mixture-of-experts models, including wide expert parallelism (WideEP). Study expert placement, load imbalance, and computation/communication trade-offs, and use the findings to guide model and system architecture choices.
  • Develop rigorous evaluationscovering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.Compare serving architectures under consistent workload conditions and GPU budgets.
  • Workwith MLE, GPU kernel, backend infrastructure, product, and customer teams toselect research priorities with measurable production impact.
  • Share results throughinternal reports, technical blogs, papers, and open-sourceartifacts, and mentorengineers and scientists on experimental design, scientific rigor, and model/systemtrade-offs.
Must-haves
  • APhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, appliedmathematics,or a closely relateddiscipline.
  • A strongpublication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.
  • Strong Python andPyTorch implementation skills, with the ability to turn ideas into experiments and working prototypes.
  • Deepknowledgeof LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-servingtrade-offs.
  • Strong experimental designskills coveringablations, baselines, metrics, statistical reasoning, and failure analysis.
  • Excellent written and verbal communication.
Nice-to-haves
  • First-author publicationsat venues such asNeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, orHPCA.
  • Experience deploying ML models or inference optimizations in production.
  • Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals.
  • Experienceapplyingpost-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation to inference quality or efficiency.
  • Open-source research artifacts, widely used benchmarks, technical blogs, or invited talksdemonstrating contributions toefficient AI systems.
  • Research or implementation experience in one or more areas of distributed inference system architecture, such as LLM request routing, PDD, multi-node serving, or MoE expert parallelism, including WideEP.

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving- Bold thinking- Constant growth- Meaningful impact- Trust and real ownership- Opportunity to shape the future of AI

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.

If you need accommodations during the application process, please let us know.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

Nebius • Amsterdam

On-site
EUR 70,000 - 90,000
Competitive compensation
Career growth opportunities
Collaborative culture
+1
Senior Software Engineer (Token Factory)
Senior Software Engineer (Token Factory)

Nebius Group • Netherlands

Hybrid
EUR 110,000 - 170,000
Competitive compensation
Career growth
Flexible work options
+2
Senior Software Engineer (Token Factory)
Senior Software Engineer (Token Factory)

Nebius • Amsterdam

Hybrid
EUR 120,000 - 155,000
Competitive compensation
International environment
Flexible work options
+2
AI Full-Stack Developer
AI Full-Stack Developer

Nebius Group • Amsterdam

On-site
EUR 80,000 - 120,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
AI Science Writer, Nebius Academy (Contract)
AI Science Writer, Nebius Academy (Contract)

Nebius • Amsterdam

On-site
EUR 65,000 - 95,000
Competitive pay
Career growth
Flexibility
+3
Senior Technical Product Manager - AI Compute Platform
Senior Technical Product Manager - AI Compute Platform

Nebius Group • Amsterdam

On-site
EUR 120,000 - 180,000
Competitive compensation
Career growth opportunities
Flexible work culture
+1
L3 Support Engineer (SOPs and Runbooks)
L3 Support Engineer (SOPs and Runbooks)

Nebius • Amsterdam

On-site
EUR 70,000 - 110,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Data Analyst (Agentic Search)
Data Analyst (Agentic Search)

Nebius Group • Amsterdam

On-site
EUR 90,000 - 140,000
Competitive pay
Career growth
Flexible work
+3
Head of Platform
Head of Platform

AI Chopping Block, Inc. • Amsterdam

On-site
EUR 140,000 - 190,000
Competitive compensation
Career growth and learning
Ownership and autonomy
+1
Customer Experience Automations Manager
Customer Experience Automations Manager

nebius • Amsterdam

On-site
EUR 120,000 - 180,000
Competitive compensation
Career growth and learningOpportunity
Flexibility and ownership
+3