Principal LLM Inference Engineer — Hybrid Lead

Hewlett Packard Enterprise Company

Spring (TX)

Hybrid

USD 160,000 - 303,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Hewlett Packard Enterprise is seeking a Principal Software Engineer to lead the model runtime within the inference platform used to operate large language models on customer-owned hardware. The role focuses on low tail latency and high GPU utilization, with architecture ownership of engine integration, batching, KV cache, and distributed execution, alongside a Kubernetes orchestration layer.

You will mentor engineers and present technical direction to leadership while collaborating with

Qualifications

  • Production experience with LLM inference engines and modification of engine internals.
  • Deep understanding of inference internals including batching, KV cache, and quantization.
  • Expertise in tensor/pipeline parallelism and GPU memory hierarchy.
  • Proficient in Kubernetes architecture, operators, CRDs, controllers and scheduling.
  • Strong programming in Go and Python; reading C++/CUDA with Nsight tooling.
  • Experience debugging multi-tier workloads (RAG, Agents, etc).
  • Excellent analytical, debugging and problem-solving abilities.

Responsibilities

  • Define and own the technical direction of the LLM serving deployment and orchestration.
  • Collaborate with performance teams to optimize time-to-first-token and tail latency.
  • Define distributed inferencing strategies across GPU memory and storage.
  • Evaluate runtimes, quantization, and serving approaches for adoption.
  • Define orchestration layer for model admission, GPU scheduling and autoscaling.
  • Mentor engineers, lead architecture reviews, and present to executives.

Skills

LLM inference engines
Inference internals
Tensor parallelism
Kubernetes architectures
Go & Python
Multi-tier debugging
Analytical ability

Education

Degree in Computer Science or related field

Job description

Hewlett Packard Enterprise is seeking a Principal Software Engineer to lead the model runtime within the inference platform used to operate large language models on customer-owned hardware. The role focuses on low tail latency and high GPU utilization, with architecture ownership of engine integration, batching, KV cache, and distributed execution, alongside a Kubernetes orchestration layer.

You will mentor engineers and present technical direction to leadership while collaborating with

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer - LLM Inference Runtime (Hybrid)
Senior Software Engineer - LLM Inference Runtime (Hybrid)

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 137,000 - 315,000
Principal LLM Inference Runtime Architect
Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 152,000 - 349,000
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Engineer – Hybrid (Remote/On-site)
Senior LLM Inference Engineer – Hybrid (Remote/On-site)

Hewlett Packard Enterprise Development LP • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 144,000 - 315,000
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Senior LLM Inference Engineer — Hybrid/Remote
Senior LLM Inference Engineer — Hybrid/Remote

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements