Senior LLM Inference Engineer — Hybrid & Impactful

HITEC

Spring, Northern (TX, KY)

Hybrid

USD 137,000 - 315,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hewlett Packard Enterprise is seeking a Senior Software Engineer to advance the model runtime within HPE AI Essentials. You'll design and implement core components of the inference platform, including engine integration, batching, KV cache management, and distributed execution, with Kubernetes orchestration support.

The role emphasizes low tail latency and high GPU utilization on customer hardware across generations.

Qualifications

  • Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving.
  • Degree in Computer Science or related field.
  • Strong debugging and profiling skills across multi‑tier workloads.

Responsibilities

  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference engineering teams to improve latency, throughput, and tail latency
  • Build and operate distributed execution capabilities across GPU memory, host memory, and storage
  • Evaluate emerging runtimes and quantization schemes and provide informed recommendations
  • Contribute to orchestration layer, including GPU scheduling and autoscaling
  • Triage customer issues end-to-end and mentor team members on engineering practices
  • Provide code and design reviews and lead by example in best practices

Skills

LLM inference
Go
Python
Kubernetes
C++/CUDA
NCCL

Education

Degree in Computer Science or related field

Tools

Nsight
Docker
KServe

Job description

Hewlett Packard Enterprise is seeking a Senior Software Engineer to advance the model runtime within HPE AI Essentials. You'll design and implement core components of the inference platform, including engine integration, batching, KV cache management, and distributed execution, with Kubernetes orchestration support.

The role emphasizes low tail latency and high GPU utilization on customer hardware across generations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer - LLM Inference Runtime (Hybrid)
Senior Software Engineer - LLM Inference Runtime (Hybrid)

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

HITEC • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer - Hybrid
Senior LLM Inference Engineer - Hybrid

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 140,000 - 315,000
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 144,000 - 315,000
Principal LLM Inference Engineer — Hybrid Lead
Principal LLM Inference Engineer — Hybrid Lead

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 160,000 - 303,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Engineer – Hybrid (Remote/On-site)
Senior LLM Inference Engineer – Hybrid (Remote/On-site)

Hewlett Packard Enterprise Development LP • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Fort Collins (CO)

On-site
USD 180,000 - 240,000
Health And Wellbeing Benefits
Personal And Professional Development
Principal LLM Inference Runtime Architect
Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 152,000 - 349,000
Senior LLM Inference Engineer — Hybrid/Remote
Senior LLM Inference Engineer — Hybrid/Remote

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000