A complete application in a minute — tailored resume and cover letter, ready to send.
Baseten powers mission-critical inference for AI companies and enables models to run faster at scale. We are seeking an Inference Performance Engineer to work across the stack from the inference engine to scheduling, serving, and routing.
You will apply techniques like prefill/decode disaggregation, speculative decoding, and KV-cache management, and you will shape model performance, latency, and cost tradeoffs for our customers.
Baseten powers mission-critical inference for AI companies and enables models to run faster at scale. We are seeking an Inference Performance Engineer to work across the stack from the inference engine to scheduling, serving, and routing.
You will apply techniques like prefill/decode disaggregation, speculative decoding, and KV-cache management, and you will shape model performance, latency, and cost tradeoffs for our customers.