Turn this role into an interview — a resume and cover letter built around what this employer wants.
Baseten is seeking an Inference Performance Engineer to accelerate the world's most demanding AI workloads. You’ll optimize the inference stack from engine to routing, applying techniques like quantization, speculative decoding, and KV-cache management to improve latency and throughput.
You’ll collaborate across teams to ship performance wins, contribute to open‑source engines, and bring up new model architectures on new hardware in a fast‑paced startup environment.
Baseten is seeking an Inference Performance Engineer to accelerate the world's most demanding AI workloads. You’ll optimize the inference stack from engine to routing, applying techniques like quantization, speculative decoding, and KV-cache management to improve latency and throughput.
You’ll collaborate across teams to ship performance wins, contribute to open‑source engines, and bring up new model architectures on new hardware in a fast‑paced startup environment.