On-Site Inference Systems Engineer — Rust/LLM Runtime

Confidential

California (MO)

On-site

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Confidential in Palo Alto is hiring a Member of Technical Staff, Inference Systems to build a from-scratch Rust runtime for LLM serving. You’ll own batching, scheduling, KV cache management and the full serving stack, shaping architecture with the founding team.

You’ll work on latency-sensitive, high-throughput systems on-site five days a week, with the opportunity to influence design decisions and optimize cost per token in a cutting-edge AI platform.

Qualifications

  • 2-10 years of experience as a backend, systems, or distributed systems engineer.
  • Deep familiarity with inference internals: attention, KV cache, batching, scheduling.
  • Hands-on time inside engines like vLLM, SGLang, or TensorRT-LLM, or experience building LLM serving infrastructure.
  • Strong systems programming skills in Rust, C++, or similar, and a genuine willingness to work in Rust day to day.
  • Experience profiling and optimizing performance‑critical systems.
  • Comfort owning an entire stack rather than a narrow slice.

Responsibilities

  • Build the serving runtime from scratch in Rust, design KV cache management and prefix caching that cut latency and cost per token.
  • Scale serving across multi-GPU and multi-node setups.
  • Profile and benchmark the full inference pipeline.
  • Work directly with the founding team on the architecture that defines the platform.

Skills

Backend systems
Distributed systems
Rust
Performance profiling

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton
NCCL

Job description

Confidential in Palo Alto is hiring a Member of Technical Staff, Inference Systems to build a from-scratch Rust runtime for LLM serving. You’ll own batching, scheduling, KV cache management and the full serving stack, shaping architecture with the founding team.

You’ll work on latency-sensitive, high-throughput systems on-site five days a week, with the opportunity to influence design decisions and optimize cost per token in a cutting-edge AI platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference Systems
Member of Technical Staff, Inference Systems

Confidential • California (MO)

On-site
USD 150,000 - 210,000
LLM Inference Engineer — High-Performance Rust Systems
LLM Inference Engineer — High-Performance Rust Systems

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
AI Inference Platform Engineer — Equity + On-Site Palo Alto
AI Inference Platform Engineer — Equity + On-Site Palo Alto

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Staff Inference Runtime Architect (Rust/Python)
Staff Inference Runtime Architect (Rust/Python)

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000