Senior Software Engineer - LLM Inference Runtime (Hybrid)

Hewlett Packard Enterprise Company

Spring (TX)

Hybrid

USD 137,000 - 315,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hewlett Packard Enterprise seeks a Senior Software Engineer to build and evolve the model runtime for the AI Essentials inference platform used by enterprises to operate large language models on customer-owned hardware, including air-gapped environments. Emphasis is on sustained execution efficiency, low tail latency, and high GPU utilization.

The role involves engine integration, batching, KV cache management, and distributed execution, with collaboration across inference teams and a Kubernetes

Qualifications

  • Experience with LLM inference runtimes or production model serving.
  • Strong understanding of batching, KV cache reuse, quantization, and speculative decoding.
  • Proficiency in Go and Python; ability to read/debug C++/CUDA.

Responsibilities

  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution.
  • Partner with inference engineering teams to improve time-to-first-token, inter-token latency, throughput per GPU, and tail latency (P95/P99).
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage.
  • Evaluate emerging runtimes and serving strategies and make recommendations on adoption.
  • Contribute to orchestration layer including model admission, GPU scheduling, cache-aware routing, and autoscaling.
  • Triage and resolve customer issues end-to-end and provide code/design reviews, mentoring the team.

Skills

LLM inference engines
Kubernetes
Go
Python
Nsight
debugging/profiling
GPU memory hierarchy
inference internals

Education

Degree in Computer Science

Tools

Nsight

Job description

Hewlett Packard Enterprise seeks a Senior Software Engineer to build and evolve the model runtime for the AI Essentials inference platform used by enterprises to operate large language models on customer-owned hardware, including air-gapped environments. Emphasis is on sustained execution efficiency, low tail latency, and high GPU utilization.

The role involves engine integration, batching, KV cache management, and distributed execution, with collaboration across inference teams and a Kubernetes

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 144,000 - 315,000
Principal LLM Inference Engineer — Hybrid Lead
Principal LLM Inference Engineer — Hybrid Lead

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 160,000 - 303,000
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer – Hybrid (Remote/On-site)
Senior LLM Inference Engineer – Hybrid (Remote/On-site)

Hewlett Packard Enterprise Development LP • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Principal LLM Inference Runtime Architect
Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 152,000 - 349,000
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer — Hybrid/Remote
Senior LLM Inference Engineer — Hybrid/Remote

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000