Senior LLM Inference Runtime Engineer (Kubernetes & GPU)

Hewlett Packard Enterprise

Durham (NC)

On-site

USD 140,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health And Wellbeing Benefits
Personal And Professional Development

Job summary

Hewlett Packard Enterprise is seeking an experienced software engineer to design, implement, and operate major components of the LLM serving runtime and its Kubernetes orchestration layer in Durham, NC. You will focus on improving latency, throughput, GPU utilization, and distributed execution while evaluating new inference technologies.

The role requires strong knowledge of inference engines, Kubernetes, Go and Python, and the ability to debug and profile C++/CUDA workloads, along with

Qualifications

  • 8+ years software engineering experience
  • 1–2+ years on LLM inference runtimes or production model serving
  • Degree in Computer Science or related field

Responsibilities

  • Design major components of LLM serving runtime and its Kubernetes orchestration layer
  • Improve latency, throughput, GPU utilization, and distributed execution
  • Evaluate emerging inference technologies and tools
  • Resolve customer issues and provide effective support
  • Contribute through code reviews, mentoring, and strong engineering practices

Skills

LLM Inference
Inference Runtime
vLLM
SGLang
TensorRT-LLM
Continuous Batching
KV Cache Management
Kubernetes
Go
Python
C++
CUDA
Tensor Parallelism
Pipeline Parallelism
NCCL
GPU Profiling

Education

Bachelor’s degree in Computer Science or related field

Job description

Hewlett Packard Enterprise is seeking an experienced software engineer to design, implement, and operate major components of the LLM serving runtime and its Kubernetes orchestration layer in Durham, NC. You will focus on improving latency, throughput, GPU utilization, and distributed execution while evaluating new inference technologies.

The role requires strong knowledge of inference engines, Kubernetes, Go and Python, and the ability to debug and profile C++/CUDA workloads, along with

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Fort Collins (CO)

On-site
USD 180,000 - 240,000
Health And Wellbeing Benefits
Personal And Professional Development
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer — Hybrid & Impactful
Senior LLM Inference Engineer — Hybrid & Impactful

HITEC • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Senior Software Engineer - LLM Inference Runtime (Hybrid)
Senior Software Engineer - LLM Inference Runtime (Hybrid)

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 137,000 - 315,000
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

HITEC • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Principal LLM Inference Runtime Architect
Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 152,000 - 349,000
Principal LLM Inference Engineer — Hybrid Lead
Principal LLM Inference Engineer — Hybrid Lead

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 160,000 - 303,000
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development