LLM Inference Engineer — High-Performance Rust Systems

Socket.dev

Palo Alto (CA)

On-site

USD 230,000 - 350,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity opportunity (0.5%)
Professional growth
High-impact work

Job summary

Premier Global Links LLC in Palo Alto, CA is seeking a Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up.

The role focuses on LLM inference, model serving, distributed systems, and inference runtime performance. Strong systems engineering skills and hands-on production experience are required; Rust experience is highly valued.

Qualifications

  • 2–10 years of experience in backend, distributed systems, or systems engineering.
  • Hands-on experience building, operating, or optimizing LLM inference or serving systems.
  • Deep understanding of transformer inference internals, including attention, KV cache, batching, and scheduling.
  • Experience with production inference engines such as vLLM, SGLang, or TensorRT-LLM.

Responsibilities

  • Build and optimize production LLM inference and model-serving systems.
  • Develop inference runtime components using Rust and other systems-level technologies.
  • Design and implement batching, scheduling, request routing, and serving infrastructure.
  • Build and optimize KV cache and prefix caching systems.
  • Scale inference workloads across multi-GPU and multi-node environments.
  • Profile, benchmark, and optimize latency, throughput, reliability, and cost.
  • Work with inference engines such as vLLM, SGLang, or TensorRT-LLM.
  • Investigate performance bottlenecks across the inference stack.
  • Contribute to core architecture and technical decisions for the platform.
  • Collaborate with a small, hands-on engineering team in a fast-paced environment.

Skills

Rust
C++
Go
Python

Education

Bachelor's or Master's in CS/CE

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton
NCCL

Job description

Premier Global Links LLC in Palo Alto, CA is seeking a Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up.

The role focuses on LLM inference, model serving, distributed systems, and inference runtime performance. Strong systems engineering skills and hands-on production experience are required; Rust experience is highly valued.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
AI Inference Platform Engineer — Equity + On-Site Palo Alto
AI Inference Platform Engineer — Equity + On-Site Palo Alto

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
On-Site Inference Systems Engineer — Rust/LLM Runtime
On-Site Inference Systems Engineer — Rust/LLM Runtime

Confidential • California (MO)

On-site
USD 150,000 - 210,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1