A complete application in a minute — tailored resume and cover letter, ready to send.
Hewlett Packard Enterprise is seeking an experienced software engineer to design, implement, and operate major components of the LLM serving runtime and its Kubernetes orchestration layer in Durham, NC. You will focus on improving latency, throughput, GPU utilization, and distributed execution while evaluating new inference technologies.
The role requires strong knowledge of inference engines, Kubernetes, Go and Python, and the ability to debug and profile C++/CUDA workloads, along with
Hewlett Packard Enterprise is seeking an experienced software engineer to design, implement, and operate major components of the LLM serving runtime and its Kubernetes orchestration layer in Durham, NC. You will focus on improving latency, throughput, GPU utilization, and distributed execution while evaluating new inference technologies.
The role requires strong knowledge of inference engines, Kubernetes, Go and Python, and the ability to debug and profile C++/CUDA workloads, along with