Senior ML Engineer: AI Inference & Performance Optimizer

Nebius

Palo Alto (CA)

Hybrid

USD 195,200 - 262,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) plan
Parental leave
Remote work reimbursement
Disability & life insurance

Job summary

Nebius in Palo Alto is seeking a Senior MLE to own end-to-end model and endpoint optimization. You will work at the intersection of distributed systems, GPU performance, and production ML engineering, debugging serving problems and delivering measurable improvements with minimal supervision.

You will deploy and optimize LLM/VLM backends, quantify latency and cost, and collaborate with researchers, kernel engineers, and platform teams to push the performance frontier of AI workloads.

Qualifications

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing LLM/VLM systems.
  • Experience with modern inference stacks such as vLLM, Triton, TensorRT-LLM, etc.
  • Strong understanding of transformer inference bottlenecks.
  • Ability to reason quantitatively about latency, throughput, cost, and quality tradeoffs.
  • Strong communication and collaboration with research, kernel, infrastructure, product, and customer teams.

Responsibilities

  • Own optimization work for specific model families, endpoints, or serving backends.
  • Run engine comparisons and recommend practical serving configurations for specific workloads.
  • Debug model quality or performance regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
  • Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems.
  • Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
  • Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
  • Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.
  • Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.
  • Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations.

Skills

Python
PyTorch
Distributed systems
LLM deployment
Model optimization
RL pipelines

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
NVIDIA Dynamo
Ray Serve
KServe

Job description

Nebius in Palo Alto is seeking a Senior MLE to own end-to-end model and endpoint optimization. You will work at the intersection of distributed systems, GPU performance, and production ML engineering, debugging serving problems and delivering measurable improvements with minimal supervision.

You will deploy and optimize LLM/VLM backends, quantify latency and cost, and collaborate with researchers, kernel engineers, and platform teams to push the performance frontier of AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead ML Systems Engineer - Large-Scale Training & RL Infra
Lead ML Systems Engineer - Large-Scale Training & RL Infra

Nebius • Palo Alto (CA)

Remote
USD 195,200 - 262,200
Health insurance
401(k) plan
Parental leave
+2
Senior ML Engineer — Training & RL Systems
Senior ML Engineer — Training & RL Systems

Socket.dev • Palo Alto (CA)

On-site
USD 195,000 - 263,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud

Nebius • United States

Remote
USD 180,000 - 250,000
Senior ML Scientist - Inference & Hardware Acceleration
Senior ML Scientist - Inference & Hardware Acceleration

Netskope • Santa Clara (CA)

On-site
USD 182,000 - 261,000
ML Inference Performance Engineer - Kernel Optimizer
ML Inference Performance Engineer - Kernel Optimizer

Cerebras Systems • United States

On-site
USD 110,000 - 150,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior ML Performance Engineer: LLM Benchmarking & GPU
Senior ML Performance Engineer: LLM Benchmarking & GPU

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
+2
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior Principal ML Inference Engineer - Edge & Embedded
Senior Principal ML Inference Engineer - Edge & Embedded

Cerence Inc. • United States

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage
Paid time off and holidays
+2
Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package