A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++. This position is an opportunity to contribute to cutting-edge AI technology in a dynamic environment.
Qualifications
Experience with building or operating ML inference or model serving systems.
Ability to reason about performance, memory usage, and system behavior under load.
Responsibilities
Design and optimize end-to-end inference pipelines.
Build inference runtimes balancing latency and throughput.
Profile and debug inference performance issues.
Skills
Strong software engineering fundamentals
Experience building ML inference systems
Performance and memory reasoning
Tools
TensorRT-LLM
vLLM
Python
C++
Job description
A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++. This position is an opportunity to contribute to cutting-edge AI technology in a dynamic environment.