An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Confidential in Palo Alto is hiring a Member of Technical Staff, Inference Systems to build a from-scratch Rust runtime for LLM serving. You’ll own batching, scheduling, KV cache management and the full serving stack, shaping architecture with the founding team.
You’ll work on latency-sensitive, high-throughput systems on-site five days a week, with the opportunity to influence design decisions and optimize cost per token in a cutting-edge AI platform.
Confidential in Palo Alto is hiring a Member of Technical Staff, Inference Systems to build a from-scratch Rust runtime for LLM serving. You’ll own batching, scheduling, KV cache management and the full serving stack, shaping architecture with the founding team.
You’ll work on latency-sensitive, high-throughput systems on-site five days a week, with the opportunity to influence design decisions and optimize cost per token in a cutting-edge AI platform.