An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Up Top seeks a hybrid IC/leader to own the technical strategy for large-scale inference performance and to build and guide a dedicated Inference Optimization team.
You will drive latency reduction, throughput improvements, and cost-per-token optimization across a modern GPU fleet, coordinating with benchmarking and routing across multiple inference engines and emerging hardware platforms.
Our client is a fast-growing consumer AI company built around privacy and user ownership. Their platform lets a large and growing base of individuals, third-party apps, and AI agents interact on a privacy-preserving foundation — with no retention of user data and no training on user inputs. The team is small, fast-moving, and product-obsessed, with a culture rooted in curiosity, ownership, and collaboration.
This is a hybrid IC / leadership role reporting to the Head of Engineering. You'll own the technical strategy for inference performance at scale, then build and lead a dedicated Inference Optimization team. The mandate is simple to state and hard to execute: drive down latency, increase throughput, and cut cost-per-token for large-scale LLM inference across a modern GPU fleet.