Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud

Hewlett Packard Enterprise

Spring (TX)

Hybrid

USD 144,000 - 315,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hewlett Packard Enterprise's Private Cloud AI organization seeks a Senior Software Engineer to advance the model runtime of HPE AI Essentials, the inference platform enabling enterprises to run large language models on owned infrastructure, including air-gapped and sovereign environments.

The role focuses on engine integration, batching, KV cache management, and distributed execution, in collaboration with inference teams and the Kubernetes stack.

Qualifications

  • Minimum of 8 years of experience in Software Engineering
  • Experience working directly on LLM inference runtimes or production model serving
  • Degree in Computer Science or related field

Responsibilities

  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption
  • Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence
  • Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team

Skills

LLM inference
Go
Python
C++/CUDA
Kubernetes
NCCL
Debugging/Profiling
Latency optimization

Education

Degree in Computer Science or related field

Tools

vLLM
SGLang
TensorRT-LLM
TGI
NVIDIA NIM

Job description

Hewlett Packard Enterprise's Private Cloud AI organization seeks a Senior Software Engineer to advance the model runtime of HPE AI Essentials, the inference platform enabling enterprises to run large language models on owned infrastructure, including air-gapped and sovereign environments.

The role focuses on engine integration, batching, KV cache management, and distributed execution, in collaboration with inference teams and the Kubernetes stack.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior Software Engineer - LLM Inference Runtime (Hybrid)
Senior Software Engineer - LLM Inference Runtime (Hybrid)

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer – Hybrid (Remote/On-site)
Senior LLM Inference Engineer – Hybrid (Remote/On-site)

Hewlett Packard Enterprise Development LP • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Engineer — Hybrid/Remote
Senior LLM Inference Engineer — Hybrid/Remote

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Principal LLM Inference Engineer — Hybrid Lead
Principal LLM Inference Engineer — Hybrid Lead

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 160,000 - 303,000
Principal LLM Inference Runtime Architect
Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 152,000 - 349,000
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
AI Infra Engineer — Hybrid, Production LLM Systems
AI Infra Engineer — Hybrid, Production LLM Systems

Hewlett Packard Enterprise Development LP • San Juan (PR), Northern (KY)

Hybrid
USD 120,000 - 180,000
Health benefits
Professional development
Inclusion and belonging