A complete application in a minute — tailored resume and cover letter, ready to send.
Hewlett Packard Enterprise is seeking an experienced software engineer to design and implement core components of the LLM runtime, focusing on engine integration, batching, and KV cache management. You will partner with distributed teams to optimize latency, throughput, and execution across deployments.
The ideal candidate has 8+ years in software engineering with 1-2+ years in LLM inference runtimes, and strong expertise in Go, Python, Kubernetes, and deep engine internals.
Design and implement key components of the LLM runtime, including engine integration, batching, and KV cache management. Partner with engineering teams to optimize inference performance, including latency, throughput, and distributed execution capabilities.
Requirements: Requires at least 8 years of software engineering experience with 1-2+ years specifically in LLM inference runtimes or production model serving. Candidates must possess strong proficiency in Go, Python, Kubernetes, and deep knowledge of inference engine internals.
Key Skills: LLM inference, Kubernetes, Go, Python, C++, CUDA, vLLM, TensorRT-LLM, Continuous batching, KV cache management, Distributed execution, GPU optimization, Tensor parallelism, Pipeline parallelism, Performance profiling
Benefits: Health and wellbeing benefits, Professional development programs, Flexible work arrangements