An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Get past ATS filters
Job summary
Hewlett Packard Enterprise in the United States (Colorado) is seeking a Senior Software Engineer to build and evolve the model runtime within HPE AI Essentials. You will design and implement components of the LLM serving deployment, tackle low tail latency, and work with a Kubernetes-based orchestration layer. The role involves optimizing for large language model workloads on customer-owned hardware, including air-gapped and sovereign environments.
Qualifications
8+ years of software engineering experience.
1–2+ years working on LLM inference runtimes or production model serving.
Degree in Computer Science or related field.
Responsibilities
Design, implement, and own major components of the LLM serving deployment and runtime.
Improve time-to-first-token, inter-token latency, and throughput per GPU.
Build and operate distributed execution capabilities across GPU memory, host memory, and RDMA storage.
Evaluate emerging runtimes and quantization schemes for adoption.
Support model admission, GPU scheduling, and autoscaling in the orchestration layer.
Triage customer issues and drive root-cause analysis to prevent recurrence.
Provide code reviews and mentorship to the team.
Skills
Go
Python
Kubernetes
C++/CUDA
NCCL
GPU memory tuning
Education
Degree in Computer Science or related field
Tools
Nsight
Job description
Hewlett Packard Enterprise in the United States (Colorado) is seeking a Senior Software Engineer to build and evolve the model runtime within HPE AI Essentials. You will design and implement components of the LLM serving deployment, tackle low tail latency, and work with a Kubernetes-based orchestration layer. The role involves optimizing for large language model workloads on customer-owned hardware, including air-gapped and sovereign environments.