Get more replies from employers
Send a job-specific resume in minutes.
Confidential in the United States seeks a senior engineer to own performance optimization for production LLM inference, focusing on latency, throughput, and cost across kernels and serving engines. You will profile GPU performance, apply quantization and batching strategies, and extend serving stacks like vLLM, TensorRT-LLM, and Triton.
You will collaborate with model and platform teams to push new architectures from works to fast, while targeting multi-GPU and accelerator-rich deployments.
Confidential in the United States seeks a senior engineer to own performance optimization for production LLM inference, focusing on latency, throughput, and cost across kernels and serving engines. You will profile GPU performance, apply quantization and batching strategies, and extend serving stacks like vLLM, TensorRT-LLM, and Triton.
You will collaborate with model and platform teams to push new architectures from works to fast, while targeting multi-GPU and accelerator-rich deployments.