A complete application in a minute — tailored resume and cover letter, ready to send.
Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco to push the performance boundaries of large-scale inference. You will collaborate with product teams to deliver end-to-end batch and online inference solutions, leveraging Ray Data and LLM engines.
The role requires familiarity with deep learning frameworks (PyTorch), distributed systems, and GPUs/CUDA, with opportunities to contribute to open source projects like vLLM and TensorRT-LLM.
Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco to push the performance boundaries of large-scale inference. You will collaborate with product teams to deliver end-to-end batch and online inference solutions, leveraging Ray Data and LLM engines.
The role requires familiarity with deep learning frameworks (PyTorch), distributed systems, and GPUs/CUDA, with opportunities to contribute to open source projects like vLLM and TensorRT-LLM.