Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise

Spring (TX)

Hybrid

USD 152,000 - 349,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hewlett Packard Enterprise is seeking a Principal Software Engineer to lead the LLM inference runtime within the AI Essentials platform. You will architect engine integration, batching strategies, KV cache reuse, and distributed execution on customer-owned hardware, while coordinating with Kubernetes orchestration and performance teams.

The role emphasizes deep expertise in LLM runtimes, multi-GPU scaling, and deployment at scale.

Qualifications

  • Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
  • Comprehensive understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
  • Tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them
  • Expert level proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
  • Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight
  • Experience with debugging/profiling multi-tier application workloads such as RAG, Agents, etc
  • Excellent analytical, debugging, and problem-solving abilities

Responsibilities

  • Define and own the technical direction of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse
  • Partner with inference performance engineering teams, with accountability for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and determine whether each runtime is adopted, developed in-house, or declined
  • Define the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences

Skills

Go language
Python language
Analytical thinking
Debugging skills
Communication

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Nsight
TensorRT-LLM
vLLM
NVIDIA NIM
TGI
C++/CUDA

Job description

Hewlett Packard Enterprise is seeking a Principal Software Engineer to lead the LLM inference runtime within the AI Essentials platform. You will architect engine integration, batching strategies, KV cache reuse, and distributed execution on customer-owned hardware, while coordinating with Kubernetes orchestration and performance teams.

The role emphasizes deep expertise in LLM runtimes, multi-GPU scaling, and deployment at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal LLM Inference Engineer — Hybrid Lead
Principal LLM Inference Engineer — Hybrid Lead

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 160,000 - 303,000
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Senior Software Engineer - LLM Inference Runtime (Hybrid)
Senior Software Engineer - LLM Inference Runtime (Hybrid)

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Engineer – Hybrid (Remote/On-site)
Senior LLM Inference Engineer – Hybrid (Remote/On-site)

Hewlett Packard Enterprise Development LP • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 144,000 - 315,000
Senior LLM Inference Engineer — Hybrid/Remote
Senior LLM Inference Engineer — Hybrid/Remote

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000