Principal Software Engineer, Inference

Hewlett Packard Enterprise

Fort Collins (CO)

Hybrid

USD 180,000 - 250,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health and wellbeing benefits
Professional development programs
Flexible work arrangements

Job summary

Hewlett Packard Enterprise in Fort Collins seeks a senior software architect to define and own the technical architecture for LLM serving, including engine integration, batching, and distributed execution strategies.

You will partner with performance engineering teams to optimize latency and throughput, mentor engineers, and present technical direction to stakeholders in a fast-paced, flexible work environment.

Qualifications

  • Requires at least 12 years of software engineering experience with a minimum of 1 year working directly on LLM inference runtimes.

Responsibilities

  • Define and own the technical architecture for LLM serving, including engine integration, batching, and distributed execution strategies.
  • Partner with performance engineering teams to optimize latency and throughput while mentoring engineers and presenting technical direction to stakeholders.

Skills

LLM inference
Distributed execution
Continuous batching
KV cache management
Quantized execution
GPU scheduling
Performance engineering
Architecture design

Tools

Kubernetes
Go
Python
C++
CUDA
TensorRT-LLM
vLLM
SGLang

Job description

Define and own the technical architecture for LLM serving, including engine integration, batching, and distributed execution strategies. Partner with performance engineering teams to optimize latency and throughput while mentoring engineers and presenting technical direction to stakeholders.

Requirements

Requires at least 12 years of software engineering experience with a minimum of 1 year working directly on LLM inference runtimes. Candidates must possess expert-level proficiency in Kubernetes, Go, Python, and deep knowledge of GPU memory hierarchies and inference internals.

Key Skills
  • LLM inference
  • vLLM
  • SGLang
  • TensorRT-LLM
  • Kubernetes
  • Go
  • Python
  • C++
  • CUDA
  • Distributed execution
  • Continuous batching
  • KV cache management
  • Quantized execution
  • GPU scheduling
  • Performance engineering
  • Architecture design
Benefits
  • Health and wellbeing benefits
  • Professional development programs
  • Flexible work arrangements
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Senior Software Engineer, Inference
Senior Software Engineer, Inference

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Staff Software Engineer, AI Inference
Staff Software Engineer, AI Inference

ChatGPT Jobs • New York (NY)

On-site
USD 180,000 - 240,000
Health Insurance
Equity
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer: LLM Inference & GPU Infra (Equity)
Staff Software Engineer: LLM Inference & GPU Infra (Equity)

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity