Senior Applied Scientist — Efficient LLM Inference & Optimization

Nebius

Palo Alto (CA)

On-site

USD 195,200 - 262,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) plan
Parental leave
Remote work stipend
Life & disability insurance

Job summary

Nebius is seeking a Senior Applied Scientist to turn frontier inference bottlenecks into production-ready solutions. You will design rigorous experiments, write high-quality code in Python and PyTorch, and collaborate with ML engineers to ship research into production.

You will own well-scoped research projects, publish credible work, mentor engineers, and define evaluation methods for latency, throughput, and cost per token across LLM/VLM inference workloads.

Qualifications

  • PhD in computer science, ML systems, or closely related field.
  • Strong publication record or equivalent artifacts in ML/efficient inference/quantization.
  • Strong coding ability in Python and PyTorch; move from idea to experiment quickly.
  • Deep understanding of LLMs, VLMs, transformer inference, decoding, quantization, and production-serving tradeoffs.
  • Strong experimental design skills, including ablations, baselines, metrics, and failure analysis.
  • Excellent written and verbal communication.

Responsibilities

  • Own focused research projects from hypothesis through experiment to production handoff.
  • Prepare internal reports, technical blogs, or papers when externally credible.
  • Partner with MLEs to ensure research prototypes become usable production components.
  • Define and execute research programs in efficient LLM/VLM inference with measurable production impact.
  • Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse and compression.
  • Build high-quality prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks; productionize them with engineers.
  • Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.
  • Publish papers, technical reports, blogs, and open-source artifacts to build external credibility for Nebius Token Factory.
  • Collaborate with MLE, GPU kernel, backend infra, product, and customer teams to select high-leverage research bets.
  • Mentor engineers and scientists on experimental design and model/system tradeoffs.

Skills

Python programming
PyTorch
Research experience
Experiment design
Communication skills

Education

PhD in CS / ML / related field

Tools

PyTorch
Triton
CUDA
TensorRT-LLM

Job description

Nebius is seeking a Senior Applied Scientist to turn frontier inference bottlenecks into production-ready solutions. You will design rigorous experiments, write high-quality code in Python and PyTorch, and collaborate with ML engineers to ship research into production.

You will own well-scoped research projects, publish credible work, mentor engineers, and define evaluation methods for latency, throughput, and cost per token across LLM/VLM inference workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Scientist: LLM/VLM Inference Expert
Senior ML Systems Scientist: LLM/VLM Inference Expert

Nebius B.V. • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Senior ML Solutions Architect – Remote Token Platform (LLM)
Senior ML Solutions Architect – Remote Token Platform (LLM)

Nebius • United States

On-site
USD 210,000 - 260,000
Health Insurance
401(k) Plan
Parental Leave
+2
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Senior ML Solutions Architect: Production LLMs & Fine-Tuning
Senior ML Solutions Architect: Production LLMs & Fine-Tuning

Nebius • United States

On-site
USD 208,000 - 261,000
Health Insurance
401(k) Plan
Parental Leave
+2
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Senior Applied Scientist, Efficient LLM Inference & Model Optimization

Nebius • Palo Alto (CA)

On-site
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
LLM/VLM Inference Optimization Research Engineer
LLM/VLM Inference Optimization Research Engineer

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+1
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior ML Engineer — Training & RL Systems
Senior ML Engineer — Training & RL Systems

Socket.dev • Palo Alto (CA)

On-site
USD 195,000 - 263,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3