Get more replies from employers
Send a job-specific resume in minutes.
Confidential in the United States seeks a senior engineer to own performance optimization for production LLM inference, focusing on latency, throughput, and cost across kernels and serving engines. You will profile GPU performance, apply quantization and batching strategies, and extend serving stacks like vLLM, TensorRT-LLM, and Triton.
You will collaborate with model and platform teams to push new architectures from works to fast, while targeting multi-GPU and accelerator-rich deployments.
Own the performance of large language models in production — the latency, the throughput, the cost-per-token. This is deep inference-optimization work: profiling and tuning at the GPU and serving-engine level to make models run faster and cheaper at scale. You'll join a small, senior team at an established enterprise software company building LLM-powered capabilities into its products.
What you'll do:
What you'll bring:
Nice to have:
A rare role where deep performance work is the whole job, not a side quest.