A leading technology firm in San Diego seeks an LLM Serving Engineer to develop scalable AI solutions. This role involves building LLM inference platforms and collaborating with teams to drive innovations in machine learning. Responsibilities include optimizing deep learning workloads and utilizing advanced techniques for efficient serving. Candidates should have strong experience with LLM packages, a solid foundation in computer science, and experience in Python development. Competitive salary and benefits are offered.
Qualifications
Hands-on experience with Triton-Inference Server and similar packages.
Strong experience in developing language models, especially using PyTorch.
Excellent understanding of algorithms and parallel programming.
Responsibilities
Build a scalable LLM inference platform using advanced techniques.
Contribute to development of LLM Serving packages.
Drive efficient serving with load balancing and routing.
Skills
Experience with LLM serving packages
Deep understanding of foundational LLMs
Experience in developing language models using PyTorch
Computer science fundamentals
Understanding of computer architecture and ML accelerators
Python development skills
Experience in optimizing deep learning workloads
Problem-solving skills
Excellent communication skills
Education
Bachelor’s degree in relevant field
Master’s degree in relevant field
PhD in relevant field
Job description
A leading technology firm in San Diego seeks an LLM Serving Engineer to develop scalable AI solutions. This role involves building LLM inference platforms and collaborating with teams to drive innovations in machine learning. Responsibilities include optimizing deep learning workloads and utilizing advanced techniques for efficient serving. Candidates should have strong experience with LLM packages, a solid foundation in computer science, and experience in Python development. Competitive salary and benefits are offered.