AI Inference Platform Engineer — Equity + On-Site Palo Alto

Premier Global Links

Palo Alto (CA)

On-site

USD 230,000 - 350,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Premier Global Links LLC is seeking an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up. You will focus on LLM inference, model serving, distributed systems, and inference runtime performance.

The role requires hands-on experience with production inference systems and strong systems engineering skills, with Rust experience highly valued. On-site in Palo Alto, CA, five days a week.

Qualifications

  • 2–10 years of backend, distributed systems, or systems engineering experience.
  • Hands-on experience building, operating, or optimizing LLM inference or serving systems.
  • Deep understanding of transformer inference internals: attention, KV cache, batching, scheduling.
  • Experience with production inference engines such as vLLM, SGLang, or TensorRT‑LLM.
  • Strong programming experience in Rust, C++, Go, or systems‑level Python/PyTorch.

Responsibilities

  • Build and optimize production LLM inference and model-serving systems.
  • Develop inference runtime components using Rust and other systems languages.
  • Design and implement batching, scheduling, request routing, and serving infrastructure.
  • Build and optimize KV cache and prefix caching systems.
  • Scale inference workloads across multi-GPU and multi-node environments.
  • Profile, benchmark, and optimize latency, throughput, reliability, and cost.
  • Work with inference engines such as vLLM, SGLang, or TensorRT‑LLM.
  • Investigate performance bottlenecks across the inference stack.
  • Contribute to core architecture and technical decisions for the platform.
  • Collaborate with a small, hands-on engineering team in a fast-paced environment.

Skills

LLM inference
Distributed systems
Rust programming
System-level Python/PyTorch
Latency optimization
KV cache
Transformer internals

Education

Bachelor's or Master's in CS/CE or related

Tools

vLLM
SGLang
TensorRT‑LLM
CUDA
Triton
NCCL

Job description

Premier Global Links LLC is seeking an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up. You will focus on LLM inference, model serving, distributed systems, and inference runtime performance.

The role requires hands-on experience with production inference systems and strong systems engineering skills, with Rust experience highly valued. On-site in Palo Alto, CA, five days a week.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Inference Platform Backend Engineer (Equity & Benefits)
Inference Platform Backend Engineer (Equity & Benefits)

Together • San Francisco (CA)

On-site
USD 160,000 - 250,000
Equity
Health insurance
Competitive compensation
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Competitive salary
Premier insurance plans
Flexible paid time off
+1
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff, Inference Systems
Member of Technical Staff, Inference Systems

Confidential • California (MO)

On-site
USD 150,000 - 210,000
On-Site Inference Systems Engineer — Rust/LLM Runtime
On-Site Inference Systems Engineer — Rust/LLM Runtime

Confidential • California (MO)

On-site
USD 150,000 - 210,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000