Staff Engineer – High-Performance Model Inference

Pantera Capital

Palo Alto (CA)

On-site

USD 180,000 - 440,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Medical coverage
Vision coverage
Dental coverage
401(k)
Disability insurance
Life insurance
Employee discounts

Job summary

xAI is building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability.

As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end, owning everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency).

Qualifications

  • Deep low-level systems programming in C/C++ or Rust.
  • Experience with large-scale, high-concurrency production serving.
  • Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
  • Strong background in system optimizations: batching, caching, load balancing, parallelism.
  • Low-level inference optimizations: GPU kernels, code generation.
  • Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics.
  • Experience with testing, benchmarking, and reliability of inference services.
  • Experience designing and implementing CI/CD infrastructure for inference.

Responsibilities

  • Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
  • Optimize latency and throughput of model inference under real production workloads.
  • Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency.
  • Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation).
  • Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels.
  • Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates.
  • Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.

Skills

Low-level systems
Large-scale serving
GPU inference engines
System optimizations
GPU kernels
Quantization
Speculative decoding
Testing & benchmarking
CI/CD for inference

Job description

xAI is building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability.

As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end, owning everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - Large-Scale Model Inference & Systems
Staff Engineer - Large-Scale Model Inference & Systems

Xai • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Staff Engineer, Inference Runtime — High-Performance AI Serving
Staff Engineer, Inference Runtime — High-Performance AI Serving

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

Visa Hunt • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
Parental leave top-up
+3
Software Engineer - Training/Inference (C++)
Software Engineer - Training/Inference (C++)

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
AI Inference Platform Engineer
AI Inference Platform Engineer

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Medical, dental and vision insurance (
Flexible PTO including Winter Break
+4
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance