Senior ML Engineer: AI Inference & Performance Optimizer

Nebius

Palo Alto (CA)

Hybrid

USD 195,200 - 262,200

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) plan
Parental leave
Remote work reimbursement
Disability & life insurance

Job summary

Nebius in Palo Alto is seeking a Senior MLE to own end-to-end model and endpoint optimization. You will work at the intersection of distributed systems, GPU performance, and production ML engineering, debugging serving problems and delivering measurable improvements with minimal supervision.

You will deploy and optimize LLM/VLM backends, quantify latency and cost, and collaborate with researchers, kernel engineers, and platform teams to push the performance frontier of AI workloads.

Qualifications

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing LLM/VLM systems.
  • Experience with modern inference stacks such as vLLM, Triton, TensorRT-LLM, etc.
  • Strong understanding of transformer inference bottlenecks.
  • Ability to reason quantitatively about latency, throughput, cost, and quality tradeoffs.
  • Strong communication and collaboration with research, kernel, infrastructure, product, and customer teams.

Responsibilities

  • Own optimization work for specific model families, endpoints, or serving backends.
  • Run engine comparisons and recommend practical serving configurations for specific workloads.
  • Debug model quality or performance regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
  • Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems.
  • Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
  • Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
  • Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.
  • Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.
  • Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations.

Skills

Python
PyTorch
Distributed systems
LLM deployment
Model optimization
RL pipelines

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
NVIDIA Dynamo
Ray Serve
KServe

Job description

Nebius in Palo Alto is seeking a Senior MLE to own end-to-end model and endpoint optimization. You will work at the intersection of distributed systems, GPU performance, and production ML engineering, debugging serving problems and delivering measurable improvements with minimal supervision.

You will deploy and optimize LLM/VLM backends, quantify latency and cost, and collaborate with researchers, kernel engineers, and platform teams to push the performance frontier of AI workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Engineer — AI Inference & Performance
Senior ML Engineer — AI Inference & Performance

Nebius Group • London (KY)

On-site
USD 160,000 - 260,000
Competitive compensation
Career growth
Flexibility
+3
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior ML Inference Engineer – High-Performance Serving
Senior ML Inference Engineer – High-Performance Serving

Amazon Inc. • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud

Nebius • United States

Remote
USD 180,000 - 250,000
Senior ML Scientist - Inference & Hardware Acceleration
Senior ML Scientist - Inference & Hardware Acceleration

Netskope • Santa Clara (CA)

On-site
USD 182,500 - 260,500
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
ML Inference & Performance Engineer
ML Inference & Performance Engineer

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
LLM Inference Optimization Engineer - Frontier Performance
LLM Inference Optimization Engineer - Frontier Performance

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000