Principal Software Engineer, Inference

Hewlett Packard Enterprise

Spring (TX)

Hybrid

USD 152,000 - 349,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Hewlett Packard Enterprise is seeking a Principal Software Engineer to lead the LLM inference runtime within the AI Essentials platform. You will architect engine integration, batching strategies, KV cache reuse, and distributed execution on customer-owned hardware, while coordinating with Kubernetes orchestration and performance teams.

The role emphasizes deep expertise in LLM runtimes, multi-GPU scaling, and deployment at scale.

Qualifications

  • Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
  • Comprehensive understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
  • Tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them
  • Expert level proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
  • Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight
  • Experience with debugging/profiling multi-tier application workloads such as RAG, Agents, etc
  • Excellent analytical, debugging, and problem-solving abilities

Responsibilities

  • Define and own the technical direction of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse
  • Partner with inference performance engineering teams, with accountability for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and determine whether each runtime is adopted, developed in-house, or declined
  • Define the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences

Skills

Go language
Python language
Analytical thinking
Debugging skills
Communication

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Nsight
TensorRT-LLM
vLLM
NVIDIA NIM
TGI
C++/CUDA

Job description

Hybrid Work Model

This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description

HPE's Private Cloud AI organization is seeking a Principal Software Engineer to lead the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The principal engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will define the architecture of that runtime – engine integration, batching, KV cache management, and distributed execution – together with the Kubernetes orchestration layer that supports it. The primary work location is as listed, but could be any other HPE site location in the US; however, remote work options will be considered.

Responsibilities
  • Define and own the technical direction of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
  • Partner with inference performance engineering teams, with accountability for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
  • Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, KV cache offload across GPU memory, host memory, and RDMA-attached storage
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and determine whether each runtime is adopted, developed in-house, or declined
  • Define the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
  • Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences
Required
  • Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
  • Comprehensive understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
  • Tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them
  • Expert level proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
  • Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight
  • Experience with debugging/profiling multi-tier application workloads such as RAG, Agents, etc
  • Excellent analytical, debugging, and problem-solving abilities
Preferred
  • Upstream contribution to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe
  • Disaggregated prefill/decode serving, or KV cache offload and reuse at scale
  • RDMA, GPUDirect Storage, InfiniBand, or RoCE
  • MIG, fractional GPU allocation, and multi-tenant GPU isolation
  • On-premises, air-gapped, or regulated enterprise software delivery
Experience and Education
  • Minimum of 12 years of experience in Software Engineering, including +1 years working directly on LLM inference runtimes or production model serving
  • Degree in Computer Science or related field
Accessibility

HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here. Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability.

What We Can Offer You
Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have — whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let’s Stay Connected

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

Job Level

TCP_05

Job

Engineering

Salary

The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.

– United States of America: Annual Salary USD 160,000 - 303,000 in Colorado // 152,000 - 349,000 in North Carolina & Texas

The listed salary range reflects base salary. Variable incentives may also be offered.

Benefits link

Information about employee benefits offered in the US can be found at https://myhperewards.com/main/new-hire-enrollment.html

Application Deadline

The estimated job application period closure is December 30 2027; this timeline is provided for transparency and internal planning purposes.

Equal Employment Opportunity

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together. Please click here: Equal Employment Opportunity.

EEO Statement

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

Legal Compliance

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hobbsnews • Spring (TX), Northern (KY)

On-site
USD 137,000 - 315,000
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hewlett Packard Enterprise Development LP • Fort Collins (CO)

On-site
USD 144,000 - 273,000
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 160,000 - 303,000
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 144,000 - 315,000
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 137,000 - 315,000
Senior Principal AI & Machine Learning Engineer, Spring, Texas, Onsite
Senior Principal AI & Machine Learning Engineer, Spring, Texas, Onsite

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 152,000 - 349,000
Health & Wellbeing
Personal & Professional Development
Unconditional Inclusion
Principal Engineer – Generative AI & LLM Platforms
Principal Engineer – Generative AI & LLM Platforms

Hewlett Packard Enterprise • San Juan (PR)

On-site
USD 180,000 - 280,000
Principal Engineer - Generative AI & LLM Platforms
Principal Engineer - Generative AI & LLM Platforms

Zerto • San Juan (PR)

Hybrid
USD 180,000 - 260,000
Senior Engineer - NPAT Content Services AI Infra.
Senior Engineer - NPAT Content Services AI Infra.

Hewlett Packard Enterprise Development LP • San Juan (PR), Northern (KY)

On-site
USD 120,000 - 180,000
Health benefits
Professional development
Inclusion and belonging