Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP

Spring, Northern (TX, KY)

Hybrid

USD 144,000 - 273,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hewlett Packard Enterprise in Spring, Texas, with hybrid work expectations, seeks a Senior Software Engineer to advance the model runtime for AI Essentials. You will work on engine integration, batching, KV cache reuse, and distributed execution, collaborating across teams to optimize latency and GPU utilization on customer-owned hardware.

The role requires deep knowledge of LLM inference, Kubernetes, and Go/Python proficiency, with opportunities to mentor teammates and contribute to code

Responsibilities

  • Design, implement, and own major components of the LLM serving deployment, including engine integration, batching, KV cache management and reuse, and quantized execution.
  • Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency.
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage.

Job description

Hewlett Packard Enterprise in Spring, Texas, with hybrid work expectations, seeks a Senior Software Engineer to advance the model runtime for AI Essentials. You will work on engine integration, batching, KV cache reuse, and distributed execution, collaborating across teams to optimize latency and GPU utilization on customer-owned hardware.

The role requires deep knowledge of LLM inference, Kubernetes, and Go/Python proficiency, with opportunities to mentor teammates and contribute to code

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer — Hybrid & GPU-Focused
Senior LLM Inference Engineer — Hybrid & GPU-Focused

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 144,000 - 315,000
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Principal AI Inference Engineer (Hybrid)
Principal AI Inference Engineer (Hybrid)

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 152,000 - 349,000
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Gen AI & LLM Platforms Engineer II (Hybrid)
Gen AI & LLM Platforms Engineer II (Hybrid)

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 90,000 - 120,000
Health & Wellbeing programs
Career development support
Inclusive culture and teams
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 144,000 - 315,000
Senior AI Platform Architect — Generative LLMs
Senior AI Platform Architect — Generative LLMs

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 180,000 - 280,000