Staff ML Systems Engineer - LLM Serving & RL

Prime Intellect

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 300,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Remote option
Visa sponsorship
Relocation support
Team off-sites
Conference attendance

Job summary

Prime Intellect is building an open frontier AI platform with a focus on scalable LLM serving and RL integration. This hybrid role spans cloud LLM serving, inference optimization, and RL systems.

You will advance our ability to evaluate and serve models trained with our RL Lab at scale, building multi-tenant serving, scheduling, and robust failover across regions. Requires 3+ years on ML/LLM infra, experience with vLLM or TensorRT-LLM, and strong Python/PyTorch skills.

Qualifications

  • 3+ years building and running large-scale ML/LLM services with clear latency/availability SLOs.
  • Hands-on experience with at least one of vLLM, SGLang, TensorRT-LLM.
  • Familiarity with distributed and disaggregated serving infra such as NVIDIA Dynamo.
  • Deep understanding of prefill vs. decode, KV-cache, batching, and speculative decoding.

Responsibilities

  • Build multi-tenant LLM serving platforms across cloud GPU fleets.
  • Design placement and scheduling for heterogeneous accelerators.
  • Implement cross-region failover and traffic shifting for resilience and cost control.
  • Develop autoscaling, routing, and load balancing to meet SLOs.
  • Integrate and optimize LLM inference frameworks (vLLM, SGLang, TensorRT-LLM).
  • Collaborate on CI/CD, observability, and incident response for ML infra.

Skills

Python
PyTorch
Kubernetes
CUDA
NCCL

Tools

vLLM
SGLang
TensorRT-LLM
NVIDIA Dynamo
Docker

Job description

Prime Intellect is building an open frontier AI platform with a focus on scalable LLM serving and RL integration. This hybrid role spans cloud LLM serving, inference optimization, and RL systems.

You will advance our ability to evaluate and serve models trained with our RL Lab at scale, building multi-tenant serving, scheduling, and robust failover across regions. Requires 3+ years on ML/LLM infra, experience with vLLM or TensorRT-LLM, and strong Python/PyTorch skills.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Remote option
Visa sponsorship
Relocation support
+2
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Equity incentives
Visa sponsorship
Relocation support
+2
Senior ML Engineer, AI Platform & Products (LLM)
Senior ML Engineer, AI Platform & Products (LLM)

BetterUp • New York (NY)

Hybrid
USD 200,000 - 275,000
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Remote ML Engineering Manager: LLM Serving & Infra
Remote ML Engineering Manager: LLM Serving & Infra

Jobgether • United States

Remote
USD 176,000 - 252,000
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallelwireless • United States

On-site
USD 140,000 - 210,000
Senior ML Serving Engineer for LLMs & Inference
Senior ML Serving Engineer for LLMs & Inference

Alldus • San Jose (CA)

On-site
USD 180,000 - 220,000
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • United States

On-site
USD 180,000 - 280,000