An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Axiōma Search is building a state-of-the-art multimodal agent inference stack in production, spanning from engine layers to serving architecture. You will help design and operate systems for low latency, high throughput, and cost efficiency while collaborating with research and engineering teams.
The role focuses on research-driven production ML and scalable infrastructure, with exposure to modern transformers and multimodal architectures in a fast-paced, VC-backed setting.
Serving a multimodal agent model in production is a different problem to serving a standard LLM. Context length, tool calls, and computer-use workloads create constraints that require co-designing the inference stack with the model team - not just bolting on a serving framework after the fact.
This is a VC-backed challenger lab building state-of-the-art computer-use agents. The inference team owns the full stack from engine layer (vLLM, SGLang) through to serving architecture (disaggregated inference, intelligent routing).
The team operates at the intersection of research and production - translating cutting-edge techniques directly into the systems behind live agent products.