A complete application in a minute — tailored resume and cover letter, ready to send.
Hewlett Packard Enterprise Development LP is seeking a Senior Software Engineer to advance the model runtime for our AI inference platform. You will design and implement core components, optimize batching and KV cache strategies, and collaborate with cross-functional teams to push performance on enterprise hardware.
You will work in a hybrid setup with occasional on-site requirements in US locations, contributing to scalable, low-latency inference in air-gapped and regulated environments.
This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office. The primary work location is as listed, but could be any other HPE site location in the US; however, remote work options will be considered.
Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.
HPE's Private Cloud AI organization is seeking a Senior Software Engineer to build and evolve the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The core engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will design and implement key components of that runtime – engine integration, batching, KV cache management, and distributed execution – together with the Kubernetes orchestration layer that supports it.
Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving. Degree in Computer Science or related field.
HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here. Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability.
We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.
We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have — whether you want to become a knowledge expert in your field or apply your skills to another division.
We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.
The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level. – United States of America: Annual Salary USD 144,000 - 273,000 in Colorado // 137,000 - 315,000 in North Carolina & Texas. The listed salary range reflects base salary. Variable incentives may also be offered.
The estimated job application period closure is December 30 2027; this timeline is provided for transparency and internal planning purposes.
Please ensure the resume you submit to us does not include any sensitive personal data. Sensitive personal data includes data revealing information about your racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, health, sex life or sexual orientation. To the extent the resume you submit does contain this type of personal data, you consent to the storing and processing of this data by HPE for the purpose of reviewing and managing your application.