Turn this role into an interview — a resume and cover letter built around what this employer wants.
Baseten is seeking an Inference Performance Engineer to accelerate AI workloads in production. You will work across the stack—from the inference engine to routing—using techniques like quantization, speculative decoding, and KV-cache management.
You’ll shape latency, throughput, and cost for customers’ models in a fast-growing startup. You’ll collaborate on open-source engines (vLLM, SGLang, TensorRT-LLM) and partner with model, infra, and customer-facing teams to ship wins.
Baseten is seeking an Inference Performance Engineer to accelerate AI workloads in production. You will work across the stack—from the inference engine to routing—using techniques like quantization, speculative decoding, and KV-cache management.
You’ll shape latency, throughput, and cost for customers’ models in a fast-growing startup. You’ll collaborate on open-source engines (vLLM, SGLang, TensorRT-LLM) and partner with model, infra, and customer-facing teams to ship wins.