A leading AI infrastructure firm is seeking a talented engineer to build and optimize a multi-tenant LLM serving platform at scale. Candidates should have over 3 years of experience in ML systems, with strong skills in tools like PyTorch and a solid understanding of inference optimization. This role offers flexible work arrangements and professional development opportunities for those passionate about democratizing AI development.
Qualifications
3+ years building and running large-scale ML/LLM services.
Hands-on with vLLM, SGLang, or TensorRT-LLM.
Deep understanding of inference mechanics and performance optimization.
Responsibilities
Build multi-tenant LLM serving platform across cloud GPU fleets.
Design scheduling algorithms for heterogeneous accelerators.
Profile kernels and optimize memory management for maximum performance.
Skills
Building ML Systems at Scale
Inference Backends
Full-Stack Debugging
Python
Kubernetes
Tools
PyTorch
CUDA
TensorRT
Job description
A leading AI infrastructure firm is seeking a talented engineer to build and optimize a multi-tenant LLM serving platform at scale. Candidates should have over 3 years of experience in ML systems, with strong skills in tools like PyTorch and a solid understanding of inference optimization. This role offers flexible work arrangements and professional development opportunities for those passionate about democratizing AI development.