Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Confidential in Palo Alto is hiring a Member of Technical Staff, Inference Systems to build a from-scratch Rust runtime for LLM serving. You’ll own batching, scheduling, KV cache management and the full serving stack, shaping architecture with the founding team.
You’ll work on latency-sensitive, high-throughput systems on-site five days a week, with the opportunity to influence design decisions and optimize cost per token in a cutting-edge AI platform.
Sponsorship: open to visa transfers (OPT, H1B) and new sponsorships (H1B, TN)
We're hiring a Member of Technical Staff, Inference Systems at a well-funded, founding-stage team building a high-performance AI inference platform from the ground up. You'll build a new inference runtime in Rust, owning batching, scheduling, request routing, and the full serving stack.
This is a from-scratch build, not a wrapper around existing tools. The team is architecting the entire runtime with latency, throughput, and cost per token as first-order concerns. Every core architectural decision is still open, and you'll be one of the people making them.
Serving LLMs at scale is a systems problem, not a model problem. Throughput and cost per token are decided by scheduling, batching, KV cache management, and how well the runtime uses GPUs across nodes. Most teams inherit these decisions from a general-purpose engine. This team is building the runtime itself, with no legacy constraints, for engineers who want to work on inference internals rather than around them.
You'll build the serving runtime from scratch in Rust, design KV cache management and prefix caching that cut latency and cost per token, scale serving across multi-GPU and multi-node setups, profile and benchmark the full inference pipeline, and work directly with the founding team on the architecture that defines the platform.
This role won’t suit you if you want remote or hybrid work, or if you prefer a defined scope with clear boundaries. The team is small, the pace is high, and you’ll be shaping architecture rather than picking up well‑specified tickets. You’ll be on‑site five days a week in Palo Alto.