Inference Performance Engineer: Accelerate LLMs & Cut Costs

US Health Partners, LLC

New York, Northern (NY, KY)

Hybrid

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Meaningful equity
US medical/dental/vision coverage for
Flexible PTO including Winter Break
Paid parental leave
Fertility and family-building stipend
401(k)
Learning and networking opportunities

Job summary

Baseten is seeking an Inference Performance Engineer to accelerate the world's most demanding AI workloads. You’ll optimize the inference stack from engine to routing, applying techniques like quantization, speculative decoding, and KV-cache management to improve latency and throughput.

You’ll collaborate across teams to ship performance wins, contribute to open‑source engines, and bring up new model architectures on new hardware in a fast‑paced startup environment.

Qualifications

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field.
  • Experience with one or more general‑purpose programming languages, such as Python or C++.
  • Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching).
  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT‑LLM.
  • Demonstrated interest and experience in LLMs.
  • Deep understanding of GPU architecture.

Responsibilities

  • Implement and productionize cutting-edge inference techniques, working deep in runtime internals. This includes quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation for structured outputs, and custom scheduling and routing algorithms.
  • Profile and optimize inference end to end, from kernel launch overhead and memory layout up to request scheduling, prefill/decode disaggregation, and cache‑aware routing.
  • Turn performance into cost savings. Improve tokens per GPU‑hour, raise utilization, and give customers and internal teams clear latency/throughput/cost tradeoffs.
  • Bring up and tune new model architectures on new hardware quickly, often in the same week they're released.
  • Build benchmarking frameworks that measure real‑world performance across model architectures, batch sizes, sequence lengths, and hardware configurations.
  • Contribute upstream to open‑source inference engines (vLLM, SGLang, TensorRT‑LLM), and partner closely with model, infrastructure, and customer‑facing teams to ship wins.

Skills

Python
C++
LLM optimization
PyTorch
TensorRT
TensorRT-LLM
LLMs
GPU architecture

Education

Bachelor's/Master's/Ph.D. in CS/Engineering/Math

Tools

CUDA
Triton
CUTLASS

Job description

Baseten is seeking an Inference Performance Engineer to accelerate the world's most demanding AI workloads. You’ll optimize the inference stack from engine to routing, applying techniques like quantization, speculative decoding, and KV-cache management to improve latency and throughput.

You’ll collaborate across teams to ship performance wins, contribute to open‑source engines, and bring up new model architectures on new hardware in a fast‑paced startup environment.

Get your free, confidential resume review.

or drag and drop your file here.