LLM Inference Performance Engineer: Speed & Efficiency

Baseten

United States

Remote

USD 150,000 - 210,000

Full time

27 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Baseten powers mission-critical inference for AI companies and enables models to run faster at scale. We are seeking an Inference Performance Engineer to work across the stack from the inference engine to scheduling, serving, and routing.

You will apply techniques like prefill/decode disaggregation, speculative decoding, and KV-cache management, and you will shape model performance, latency, and cost tradeoffs for our customers.

Qualifications

  • Experience with inference optimization techniques and end-to-end profiling.

Responsibilities

  • Implement and productionize inference techniques in runtime internals.

Skills

Inference optimization
Runtime internals
Performance profiling
Kernel optimization

Education

Bachelor's / Master's / PhD in CS, Engineering or Math

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Baseten powers mission-critical inference for AI companies and enables models to run faster at scale. We are seeking an Inference Performance Engineer to work across the stack from the inference engine to scheduling, serving, and routing.

You will apply techniques like prefill/decode disaggregation, speculative decoding, and KV-cache management, and you will shape model performance, latency, and cost tradeoffs for our customers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Inference Performance Engineer
LLM Inference Performance Engineer

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break
Software Engineer- Inference Performance
Software Engineer- Inference Performance

Baseten • United States

Remote
USD 150,000 - 210,000
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4