LLM Inference Performance Engineer

Emploive

New York (NY)

On-site

USD 180,000 - 360,000

Full time

21 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Insurance for dependents
Winter Break

Job summary

Baseten is seeking an Inference Performance Engineer to accelerate AI workloads in production. You will work across the stack—from the inference engine to routing—using techniques like quantization, speculative decoding, and KV-cache management.

You’ll shape latency, throughput, and cost for customers’ models in a fast-growing startup. You’ll collaborate on open-source engines (vLLM, SGLang, TensorRT-LLM) and partner with model, infra, and customer-facing teams to ship wins.

Qualifications

  • Bachelor's, Master's, or Ph.D. in CS, Engineering, Mathematics, or related field.
  • Experience with Python or C++ and ML libraries (PyTorch, TensorRT).
  • Familiarity with LLM optimization techniques and GPU architecture.

Responsibilities

  • Implement and productionize inference techniques in runtime internals (quantization, speculative decoding, KV-cache).
  • Profile and optimize end-to-end inference, from kernel launches to routing and batching.
  • Build benchmarking frameworks and contribute upstream to open-source inference engines.

Skills

Python
C++
LLM optimization
GPU architecture
PyTorch
TensorRT

Education

Bachelor's degree
Master's degree
PhD

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Baseten is seeking an Inference Performance Engineer to accelerate AI workloads in production. You will work across the stack—from the inference engine to routing—using techniques like quantization, speculative decoding, and KV-cache management.

You’ll shape latency, throughput, and cost for customers’ models in a fast-growing startup. You’ll collaborate on open-source engines (vLLM, SGLang, TensorRT-LLM) and partner with model, infra, and customer-facing teams to ship wins.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Inference Performance Engineer: Speed & Efficiency
LLM Inference Performance Engineer: Speed & Efficiency

Baseten • United States

Remote
USD 150,000 - 210,000
Software Engineer- Inference Performance
Software Engineer- Inference Performance

Baseten • United States

Remote
USD 150,000 - 210,000
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
LLM Inference Engineer (Mid, Senior, Staff)
LLM Inference Engineer (Mid, Senior, Staff)

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Software Engineer- Inference Performance Baseten · New York City, NY Full-time · Hybrid $180,000–360,000 25 minutes ago
Software Engineer- Inference Performance Baseten · New York City, NY Full-time · Hybrid $180,000–360,000 25 minutes ago

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break