Principal Software Engineer, Inference

Hewlett Packard Enterprise

Durham (NC)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health & Wellbeing
Personal & Professional Development

Job summary

Hewlett Packard Enterprise seeks a senior architect to define and own the technical architecture for LLM inference runtimes, focusing on execution efficiency and distributed inferencing strategies on customer-owned hardware. You will lead design reviews and mentor engineers.

You will collaborate with performance teams to optimize latency and throughput, apply continuous batching, and advance GPU memory management and inference internals in a fast-paced AI systems environment.

Qualifications

  • Requires at least 12 years of software engineering experience with a minimum of 1 year working directly on LLM inference runtimes.
  • Candidates must possess expert-level proficiency in Kubernetes, Python, Go, and deep knowledge of GPU memory hierarchies and inference internals.

Responsibilities

  • Define and own the technical architecture for LLM inference runtimes.
  • Lead design reviews, mentor engineers, and partner with performance teams to optimize latency and throughput on customer-owned hardware.

Skills

LLM Inference
Kubernetes
Python
Go
C++
CUDA
TensorRT-LLM
vLLM
Distributed Systems
GPU Optimization
Continuous Batching
KV Cache Management
Architecture Design
Performance Engineering
Machine Learning

Job description

Define and own the technical architecture for LLM inference runtimes, focusing on execution efficiency and distributed inferencing strategies. Lead design reviews, mentor engineers, and partner with performance teams to optimize latency and throughput on customer-owned hardware.

Requirements

Requires at least 12 years of software engineering experience with a minimum of 1 year working directly on LLM inference runtimes. Candidates must possess expert-level proficiency in Kubernetes, Python, Go, and deep knowledge of GPU memory hierarchies and inference internals.

Key Skills
  • LLM Inference
  • Kubernetes
  • Python
  • Go
  • C++
  • CUDA
  • TensorRT-LLM
  • vLLM
  • Distributed Systems
  • GPU Optimization
  • Continuous Batching
  • KV Cache Management
  • Architecture Design
  • Performance Engineering
  • Machine Learning
Benefits
  • Health & Wellbeing
  • Personal & Professional Development
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior Software Engineer - AI Inference
Senior Software Engineer - AI Inference

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
Staff Software Engineer, AI Inference
Staff Software Engineer, AI Inference

ChatGPT Jobs • New York (NY)

On-site
USD 180,000 - 240,000
Health Insurance
Equity
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000