Staff Engineer - LLM Serving & Inference at Scale

Primeintellect

San Francisco (CA)

Hybrid

USD 150,000 - 300,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote or SF office
Visa sponsorship
Relocation support
Professional development budget
Team off-sites
Equity incentives

Job summary

Prime Intellect in San Francisco is seeking an experienced ML systems engineer to scale LLM serving and RL components across cloud GPU fleets.

You will build multi-tenant serving platforms, optimize inference and integrate frameworks like vLLM, SGLang, and TensorRT-LLM, while implementing robust CI/CD and observability.

This hybrid role offers remote or San Francisco office options, visa sponsorship, relocation support, and meaningful equity to shape frontier AI infrastructure.

Qualifications

  • 3+ years building and running large-scale ML/LLM services with clear latency/availability SLOs.
  • Hands-on with at least one of vLLM, SGLang, TensorRT-LLM.
  • Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.
  • Deep understanding of prefill vs. decode, KV-cache, batching, sampling, speculative decoding, parallelism strategies.
  • Full-Stack debugging across CUDA/NCCL, drivers/kernels, containers, service mesh and storage.

Responsibilities

  • Build a multi-tenant LLM serving platform across cloud GPU fleets.
  • Design placement and scheduling for heterogeneous accelerators.
  • Implement multi-region/zone failover and traffic shifting for resilience and cost control.
  • Develop autoscaling, routing, and load balancing for throughput and latency SLOs.
  • Optimize model distribution and cold-start times across clusters.
  • Integrate and contribute to LLM inference frameworks like vLLM, SGLang, TensorRT-LLM.
  • Tune configurations for tensor/pipeline/expert parallelism and memory management.
  • Profile kernels and memory bandwidth; apply quantization and speculative decoding.
  • Embed and optimize distributed inference within our RL stack.
  • Establish CI/CD with reproducible builds and performance gates.
  • Build observability via metrics, logs, tracing; incident response and SLO management.

Skills

ML systems at scale
Inference backends
Distributed serving
Inference internals
Full-stack debugging

Tools

vLLM
SGLang
TensorRT-LLM
NVIDIA Dynamo

Job description

Prime Intellect in San Francisco is seeking an experienced ML systems engineer to scale LLM serving and RL components across cloud GPU fleets.

You will build multi-tenant serving platforms, optimize inference and integrate frameworks like vLLM, SGLang, and TensorRT-LLM, while implementing robust CI/CD and observability.

This hybrid role offers remote or San Francisco office options, visa sponsorship, relocation support, and meaningful equity to shape frontier AI infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Systems Engineer - LLM Serving & RL
Staff ML Systems Engineer - LLM Serving & RL

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Remote option
Visa sponsorship
Relocation support
+2
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Remote option
Visa sponsorship
Relocation support
+2
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Equity incentives
Visa sponsorship
Relocation support
+2
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Primeintellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Remote or SF office
Visa sponsorship
Relocation support
+3
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Staff Software Engineer: LLM Inference & GPU Infra (Equity)
Staff Software Engineer: LLM Inference & GPU Infra (Equity)

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity