Gimlet Labs, Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference systems under real production constraints, ensuring fast, predictable, and scalable performance. Key responsibilities include optimizing inference pipelines and managing cache allocation. Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance tuning capabilities. Familiarity with Python and C++ is preferred.
Qualifications
Strong software engineering fundamentals.
Experience building or operating ML inference or model serving systems.
Comfort reasoning about performance, memory usage, and system behavior under load.
Responsibilities
Design and optimize end‑to‑end inference pipelines from request ingestion through execution and response.
Build and evolve inference runtimes that balance latency, throughput, and concurrency under real‑world load.
Manage KV cache allocation, placement, reuse, and eviction across models and requests.
Skills
Software engineering fundamentals
Building ML inference systems
Performance reasoning
Tools
TensorRT-LLM
vLLM
Python
C++
Job description
Gimlet Labs, Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference systems under real production constraints, ensuring fast, predictable, and scalable performance. Key responsibilities include optimizing inference pipelines and managing cache allocation. Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance tuning capabilities. Familiarity with Python and C++ is preferred.