Senior LLM Inference Engineer – Hybrid (Remote/On-site)

Hewlett Packard Enterprise Development LP

Fort Collins (CO)

Hybrid

USD 144,000 - 273,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Hewlett Packard Enterprise Development LP is seeking a Senior Software Engineer to advance the model runtime for our AI inference platform. You will design and implement core components, optimize batching and KV cache strategies, and collaborate with cross-functional teams to push performance on enterprise hardware.

You will work in a hybrid setup with occasional on-site requirements in US locations, contributing to scalable, low-latency inference in air-gapped and regulated environments.

Qualifications

  • Familiar with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals.
  • Strong understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
  • Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them.
  • Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling.
  • Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight.
  • Familiar with debugging/profiling multi-tier application workloads such as RAG, Agents.
  • Excellent analytical, debugging, and problem-solving abilities.

Responsibilities

  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption
  • Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence
  • Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team

Skills

LLM inference engines
inference internals
Kubernetes platforms
Go programming
Python programming
C++/CUDA

Education

Bachelor's degree in Computer Science or related field

Tools

Nsight
NVIDIA NIM
Kubernetes

Job description

Hewlett Packard Enterprise Development LP is seeking a Senior Software Engineer to advance the model runtime for our AI inference platform. You will design and implement core components, optimize batching and KV cache strategies, and collaborate with cross-functional teams to push performance on enterprise hardware.

You will work in a hybrid setup with occasional on-site requirements in US locations, contributing to scalable, low-latency inference in air-gapped and regulated environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer - LLM Inference Runtime (Hybrid)
Senior Software Engineer - LLM Inference Runtime (Hybrid)

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud
Senior LLM Inference Engineer - Hybrid, Edge-to-Cloud

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 144,000 - 315,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Senior LLM Inference Engineer — Hybrid/Remote
Senior LLM Inference Engineer — Hybrid/Remote

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Principal LLM Inference Engineer — Hybrid Lead
Principal LLM Inference Engineer — Hybrid Lead

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 160,000 - 303,000
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Principal LLM Inference Runtime Architect
Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 152,000 - 349,000
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000