LLM Inference Performance Engineer

Xapply

San Francisco (CA)

On-site

USD 170,000 - 230,000

Full time

16 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation and equity
Medical, dental, vision insurance (US)
Flexible PTO including Winter Break
Parental leave
Carrot stipend for family-building
401(k) (US)
Exposure to ML startups

Job summary

Baseten is seeking an Software Engineer focusing on inference performance to accelerate AI workloads at scale. You will optimize runtime internals, quantization, decoding, and KV-cache management across GPUs, improving speed and cost efficiency for customers.

You will work on cutting-edge baseten inference stack, benchmark across architectures, and contribute to open-source engines. A strong background in GPU/LLM optimization is essential.

Qualifications

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field.
  • Experience with Python or C++ programming languages.
  • Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching).
  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
  • Demonstrated interest and experience in LLMs.
  • Deep understanding of GPU architecture.

Responsibilities

  • Implement and productionize cutting-edge inference techniques, working deep in runtime internals.
  • Profile and optimize inference end to end, from kernel launch to routing and cache-aware scheduling.
  • Turn performance into cost savings by improving tokens per GPU-hour, utilization, and latency/throughput tradeoffs.
  • Bring up and tune new model architectures on new hardware quickly.
  • Build benchmarking frameworks that measure real-world performance across model architectures and hardware.
  • Contribute upstream to open-source inference engines (vLLM, SGLang, TensorRT-LLM) and collaborate with teams to ship wins.

Skills

Python
C++
PyTorch
TensorRT
TensorRT-LLM
Quantization
Speculative Decoding
Continuous Batching
GPU Architecture
CUDA
Triton
CUTLASS

Education

Bachelor's/Master's/Ph.D. in CS/Engineering/Math

Tools

TensorRT
Triton
CUTLASS
vLLM
SGLang
TensorRT-LLM

Job description

Baseten is seeking an Software Engineer focusing on inference performance to accelerate AI workloads at scale. You will optimize runtime internals, quantization, decoding, and KV-cache management across GPUs, improving speed and cost efficiency for customers.

You will work on cutting-edge baseten inference stack, benchmark across architectures, and contribute to open-source engines. A strong background in GPU/LLM optimization is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Inference Performance Engineer
LLM Inference Performance Engineer

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break
LLM Inference Performance Engineer: Speed & Efficiency
LLM Inference Performance Engineer: Speed & Efficiency

Baseten • United States

Remote
USD 150,000 - 210,000
Software Engineer- Inference Performance
Software Engineer- Inference Performance

Baseten • United States

Remote
USD 150,000 - 210,000
Senior Software Engineer, LLM Performance Tooling
Senior Software Engineer, LLM Performance Tooling

BaseTen • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Equity
Health plans for dependents
Flexible PTO (Winter Break)
+3
LLM Inference Optimization Engineer - Frontier Performance
LLM Inference Optimization Engineer - Frontier Performance

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000
LLM Performance Engineer — GPU & HPC
LLM Performance Engineer — GPU & HPC

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Software Engineer- Inference Performance
Software Engineer- Inference Performance

Baseten • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Meaningful equity
Medical, dental, vision insurance (US)
+4
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Equity
Benefits
Software Engineer- Inference Performance Baseten · New York City, NY Full-time · Hybrid $180,000–360,000 25 minutes ago
Software Engineer- Inference Performance Baseten · New York City, NY Full-time · Hybrid $180,000–360,000 25 minutes ago

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package