Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise

Fort Collins (CO)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health And Wellbeing Benefits
Personal And Professional Development

Job summary

Hewlett Packard Enterprise seeks an experienced software engineer to design, implement, and operate core components of an enterprise LLM serving runtime. You will drive latency/throughput improvements, contribute to Kubernetes orchestration, and mentor teammates across reviews and technical guidance.

The role requires deep knowledge of inference engines and internals, plus strong Go, Python, C++, and CUDA debugging skills.

Qualifications

  • Requires at least eight years of software engineering experience.
  • One to two+ years working directly on LLM inference runtimes or production model serving.
  • Strong knowledge of inference engines, Kubernetes architecture, Go and Python; debugging and profiling C++/CUDA workloads.

Responsibilities

  • Design, implement, and operate major components of an enterprise LLM serving runtime.
  • Improve inference latency and throughput.
  • Contribute to Kubernetes orchestration, resolve customer issues, and provide technical leadership through reviews and mentoring.

Skills

LLM Inference
Inference Runtime Engineering
vLLM
SGLang
TensorRT-LLM
Continuous Batching
KV Cache Management
Quantization
Speculative Decoding
Tensor And Pipeline Parallelism
NCCL
Kubernetes
Go
Python
C++
CUDA

Education

CS or related degree

Tools

Kubernetes
Go
Python
C++
CUDA

Job description

Hewlett Packard Enterprise seeks an experienced software engineer to design, implement, and operate core components of an enterprise LLM serving runtime. You will drive latency/throughput improvements, contribute to Kubernetes orchestration, and mentor teammates across reviews and technical guidance.

The role requires deep knowledge of inference engines and internals, plus strong Go, Python, C++, and CUDA debugging skills.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Runtime Engineer (Kubernetes & GPU)
Senior LLM Inference Runtime Engineer (Kubernetes & GPU)

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 140,000 - 180,000
Health And Wellbeing Benefits
Personal And Professional Development
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

HITEC • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Principal LLM Inference Runtime Architect
Principal LLM Inference Runtime Architect

Hewlett Packard Enterprise • Spring (TX)

Hybrid
USD 152,000 - 349,000
Senior Software Engineer - LLM Inference Runtime (Hybrid)
Senior Software Engineer - LLM Inference Runtime (Hybrid)

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer — Hybrid & Impactful
Senior LLM Inference Engineer — Hybrid & Impactful

HITEC • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
Principal LLM Inference Engineer — Hybrid Lead
Principal LLM Inference Engineer — Hybrid Lead

Hewlett Packard Enterprise Company • Spring (TX)

Hybrid
USD 160,000 - 303,000
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer - Hybrid
Senior LLM Inference Engineer - Hybrid

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 140,000 - 315,000