Senior LLM Inference & Serving Engineer

Premier Global Links LLC

Palo Alto, Northern (CA, KY)

Hybrid

USD 230,000 - 350,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity 0.5%

Job summary

Premier Global Links LLC is seeking an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up. This role focuses on LLM inference, model serving, distributed systems, and inference runtime performance in Palo Alto.

You will work with Rust and other systems-level technologies to design batching, routing, and caching, and to scale workloads across multi-GPU and multi-node environments.

Qualifications

  • 2–10 years of backend or systems engineering experience.
  • Hands-on experience with LLM inference or serving systems.
  • Deep understanding of transformer inference internals (attention, KV cache, batching).
  • Experience with production inference engines (vLLM, SGLang, TensorRT-LLM).
  • Strong Rust, C++, Go, or systems-level Python/PyTorch for systems work.

Responsibilities

  • Build and optimize production LLM inference and model-serving systems.
  • Develop inference runtime components using Rust and other systems-level technologies.
  • Design and implement batching, scheduling, request routing, and serving infrastructure.
  • Build and optimize KV cache and prefix caching systems.
  • Scale inference workloads across multi-GPU and multi-node environments.
  • Profile, benchmark, and optimize latency, throughput, reliability, and cost.
  • Work with inference engines such as vLLM, SGLang, or TensorRT-LLM.
  • Investigate performance bottlenecks across the inference stack.
  • Contribute to core architecture and technical decisions for the platform.
  • Collaborate with a small, hands-on engineering team in a fast-paced environment.

Skills

LLM inference
Distributed systems
Rust
Model serving
Performance optimization
Python/PyTorch (systems-level)

Education

Bachelor's or Master's in CS/CE
Relevant graduate degree preferred

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton

Job description

Premier Global Links LLC is seeking an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up. This role focuses on LLM inference, model serving, distributed systems, and inference runtime performance in Palo Alto.

You will work with Rust and other systems-level technologies to design batching, routing, and caching, and to scale workloads across multi-GPU and multi-node environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Engineer — High-Performance Rust Systems
LLM Inference Engineer — High-Performance Rust Systems

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
AI Inference Platform Engineer — Equity + On-Site Palo Alto
AI Inference Platform Engineer — Equity + On-Site Palo Alto

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior AI Systems Engineer: LLM Inference & Optimization
Senior AI Systems Engineer: LLM Inference & Optimization

Showcify • United States

Remote
USD 180,000 - 240,000