Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
ByteDance's Inference Infrastructure team is seeking engineers to design, build, and operate cloud-native GPU-accelerated ML infrastructure at scale, including vLLM, SGLang, and TensorRT-LLM work.
You will join a world-class team within Core Compute Infrastructure, contributing to open-source ecosystems, scheduling, and orchestration across multi-cloud environments. A PhD and deep expertise in distributed systems help you drive high-performance inference at scale.
ByteDance's Inference Infrastructure team is seeking engineers to design, build, and operate cloud-native GPU-accelerated ML infrastructure at scale, including vLLM, SGLang, and TensorRT-LLM work.
You will join a world-class team within Core Compute Infrastructure, contributing to open-source ecosystems, scheduling, and orchestration across multi-cloud environments. A PhD and deep expertise in distributed systems help you drive high-performance inference at scale.