Get more replies from employers
Send a job-specific resume in minutes.
ByteDance in San Jose seeks a Senior GPU Inference Engineer to drive the iteration of large model inference engines and optimize GPU memory, latency, and throughput across multi-card systems.
You will design distributed parallel solutions (TP/PP/sequence/MoE) and push performance across vLLM, TensorRT-LLM, and related frameworks while collaborating with cross-functional teams to scale cutting-edge AI capabilities for internal products.
ByteDance in San Jose seeks a Senior GPU Inference Engineer to drive the iteration of large model inference engines and optimize GPU memory, latency, and throughput across multi-card systems.
You will design distributed parallel solutions (TP/PP/sequence/MoE) and push performance across vLLM, TensorRT-LLM, and related frameworks while collaborating with cross-functional teams to scale cutting-edge AI capabilities for internal products.