Get more replies from employers
Send a job-specific resume in minutes.
ByteDance Seed Infra in San Jose is seeking a Research Engineer to design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs. You will work on inference engines, serving frameworks, and end-to-end deployment pipelines, aiming to reduce latency and increase throughput.
Ideal candidates will have strong C/C++ and Python skills, hands-on ML framework experience (PyTorch/TensorFlow), and background in GPU-accelerated optimization.
ByteDance Seed Infra in San Jose is seeking a Research Engineer to design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs. You will work on inference engines, serving frameworks, and end-to-end deployment pipelines, aiming to reduce latency and increase throughput.
Ideal candidates will have strong C/C++ and Python skills, hands-on ML framework experience (PyTorch/TensorFlow), and background in GPU-accelerated optimization.