Senior LLM Inference Engineer — Hybrid & GPU-Focused

Hewlett Packard Enterprise Development LP

Spring (TX)

Hybrid

USD 144,000 - 315,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hewlett Packard Enterprise Development LP in the United States seeks a Senior Software Engineer to design and evolve the model runtime for the AI inference platform, focusing on low tail latency and high GPU utilization on customer-owned hardware. You will implement core components, work with inference teams, and enhance batching, KV cache management, and distributed execution across GPUs.

The role is hybrid with an expected 2 days per week from an HPE office, with remote options considered.

Qualifications

  • Minimum of 8 years of experience in Software Engineering.
  • Degree in Computer Science or related field.
  • Experience with LLM inference runtimes or production model serving is required.

Responsibilities

  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption
  • Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence
  • Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team

Skills

LLM engines
Batching
KV cache
GPU memory
Kubernetes
Go
Python
C++/CUDA
Profiling
Nsight

Education

Bachelor's degree in Computer Science

Tools

Nsight

Job description

Hewlett Packard Enterprise Development LP in the United States seeks a Senior Software Engineer to design and evolve the model runtime for the AI inference platform, focusing on low tail latency and high GPU utilization on customer-owned hardware. You will implement core components, work with inference teams, and enhance batching, KV cache management, and distributed execution across GPUs.

The role is hybrid with an expected 2 days per week from an HPE office, with remote options considered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Principal AI Inference Engineer (Hybrid)
Principal AI Inference Engineer (Hybrid)

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 152,000 - 349,000
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Gen AI & LLM Platforms Engineer II (Hybrid)
Gen AI & LLM Platforms Engineer II (Hybrid)

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 90,000 - 120,000
Health & Wellbeing programs
Career development support
Inclusive culture and teams
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 144,000 - 315,000
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 152,000 - 349,000
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements