Member of Technical Staff, Inference Systems

Confidential

California (MO)

On-site

USD 150,000 - 210,000

Full time

12 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Confidential in Palo Alto is hiring a Member of Technical Staff, Inference Systems to build a from-scratch Rust runtime for LLM serving. You’ll own batching, scheduling, KV cache management and the full serving stack, shaping architecture with the founding team.

You’ll work on latency-sensitive, high-throughput systems on-site five days a week, with the opportunity to influence design decisions and optimize cost per token in a cutting-edge AI platform.

Qualifications

  • 2-10 years of experience as a backend, systems, or distributed systems engineer.
  • Deep familiarity with inference internals: attention, KV cache, batching, scheduling.
  • Hands-on time inside engines like vLLM, SGLang, or TensorRT-LLM, or experience building LLM serving infrastructure.
  • Strong systems programming skills in Rust, C++, or similar, and a genuine willingness to work in Rust day to day.
  • Experience profiling and optimizing performance‑critical systems.
  • Comfort owning an entire stack rather than a narrow slice.

Responsibilities

  • Build the serving runtime from scratch in Rust, design KV cache management and prefix caching that cut latency and cost per token.
  • Scale serving across multi-GPU and multi-node setups.
  • Profile and benchmark the full inference pipeline.
  • Work directly with the founding team on the architecture that defines the platform.

Skills

Backend systems
Distributed systems
Rust
Performance profiling

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton
NCCL

Job description

Member of Technical Staff, Inference Systems (On-site)

Sponsorship: open to visa transfers (OPT, H1B) and new sponsorships (H1B, TN)

We're hiring a Member of Technical Staff, Inference Systems at a well-funded, founding-stage team building a high-performance AI inference platform from the ground up. You'll build a new inference runtime in Rust, owning batching, scheduling, request routing, and the full serving stack.

This is a from-scratch build, not a wrapper around existing tools. The team is architecting the entire runtime with latency, throughput, and cost per token as first-order concerns. Every core architectural decision is still open, and you'll be one of the people making them.

The problem you'd help solve:

Serving LLMs at scale is a systems problem, not a model problem. Throughput and cost per token are decided by scheduling, batching, KV cache management, and how well the runtime uses GPUs across nodes. Most teams inherit these decisions from a general-purpose engine. This team is building the runtime itself, with no legacy constraints, for engineers who want to work on inference internals rather than around them.

You'll build the serving runtime from scratch in Rust, design KV cache management and prefix caching that cut latency and cost per token, scale serving across multi-GPU and multi-node setups, profile and benchmark the full inference pipeline, and work directly with the founding team on the architecture that defines the platform.

You’ll likely be a fit if you have:
  • 2-10 years of experience as a backend, systems, or distributed systems engineer
  • Deep familiarity with inference internals: attention, KV cache, batching, scheduling
  • Hands‑on time inside engines like vLLM, SGLang, or TensorRT‑LLM, or experience building LLM serving infrastructure
  • Strong systems programming skills in Rust, C++, or similar, and a genuine willingness to work in Rust day to day
  • Experience profiling and optimizing performance‑critical systems
  • Comfort owning an entire stack rather than a narrow slice
Nice to have:
  • Multi‑GPU and multi‑node serving experience
  • CUDA, Triton, or NCCL experience
  • Production Rust
  • Early‑stage startup experience
What you won’t find here:

This role won’t suit you if you want remote or hybrid work, or if you prefer a defined scope with clear boundaries. The team is small, the pace is high, and you’ll be shaping architecture rather than picking up well‑specified tickets. You’ll be on‑site five days a week in Palo Alto.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff — Inference
Member of Technical Staff — Inference

RadixArk • Palo Alto (CA)

On-site
USD 190,000 - 260,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000