Principal AI Inference Engineer (Hybrid)

Hewlett Packard Enterprise Development LP

Spring (TX)

Hybrid

USD 152,000 - 349,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hewlett Packard Enterprise Development LP is seeking a Principal Software Engineer to lead the model runtime for the AI inference platform, focusing on low tail latency and high GPU utilization. You will drive architecture decisions, batching, and distributed execution across environments.

The role requires guiding a team of engineers, coordinating with performance groups, and leveraging Kubernetes to deploy and optimize the stack on-prem and cloud. Hybrid work is available.

Qualifications

  • Minimum 12 years of software engineering experience, including +1 year on LLM inference runtimes.
  • Degree in Computer Science or related field.
  • Strong Go, Python, C++/CUDA development and debugging skills.

Responsibilities

  • Define architecture of the model runtime, engine integration and distributed execution.
  • Mentor engineers, lead design reviews and present directions to stakeholders.
  • Collaborate with performance teams to optimize latency and GPU utilization.

Skills

LLM inference
Go
Python
C++
CUDA

Education

Bachelor's degree in Computer Science

Tools

Kubernetes
Nsight
TensorRT-LLM

Job description

Hewlett Packard Enterprise Development LP is seeking a Principal Software Engineer to lead the model runtime for the AI inference platform, focusing on low tail latency and high GPU utilization. You will drive architecture decisions, batching, and distributed execution across environments.

The role requires guiding a team of engineers, coordinating with performance groups, and leveraging Kubernetes to deploy and optimize the stack on-prem and cloud. Hybrid work is available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer — Hybrid & GPU-Focused
Senior LLM Inference Engineer — Hybrid & GPU-Focused

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 144,000 - 315,000
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
LLM Inference Runtime Architect | Performance & Scalability
LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Senior AI Inference Runtime Architect
Senior AI Inference Runtime Architect

Arm • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Principal AI Inference Cloud Architect
Principal AI Inference Cloud Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid Principal Engineer: Generative AI & LLM Platforms
Hybrid Principal Engineer: Generative AI & LLM Platforms

Hewlett Packard Enterprise Development LP • San Juan (PR)

Hybrid
USD 120,000 - 190,000
Senior AI Architect - Private Cloud & Agentic AI
Senior AI Architect - Private Cloud & Agentic AI

Socket.dev • San Jose (CA)

Hybrid
USD 194,000 - 413,000
Hybrid AI & Analytics Engineer (Finance)
Hybrid AI & Analytics Engineer (Finance)

Hewlett Packard Enterprise Company in • Fort Collins (CO)

Hybrid
USD 76,000 - 144,000
Hybrid Senior Applied ML Engineer: Deploy AI at Scale
Hybrid Senior Applied ML Engineer: Deploy AI at Scale

Hewlett Packard Enterprise • San Jose (CA)

Hybrid
USD 155,000 - 315,000
Hybrid work