LLM Inference Runtime Architect | Performance & Scalability

Hewlett Packard Enterprise

Durham (NC)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health & Wellbeing
Personal & Professional Development

Job summary

Hewlett Packard Enterprise seeks a senior architect to define and own the technical architecture for LLM inference runtimes, focusing on execution efficiency and distributed inferencing strategies on customer-owned hardware. You will lead design reviews and mentor engineers.

You will collaborate with performance teams to optimize latency and throughput, apply continuous batching, and advance GPU memory management and inference internals in a fast-paced AI systems environment.

Qualifications

  • Requires at least 12 years of software engineering experience with a minimum of 1 year working directly on LLM inference runtimes.
  • Candidates must possess expert-level proficiency in Kubernetes, Python, Go, and deep knowledge of GPU memory hierarchies and inference internals.

Responsibilities

  • Define and own the technical architecture for LLM inference runtimes.
  • Lead design reviews, mentor engineers, and partner with performance teams to optimize latency and throughput on customer-owned hardware.

Skills

LLM Inference
Kubernetes
Python
Go
C++
CUDA
TensorRT-LLM
vLLM
Distributed Systems
GPU Optimization
Continuous Batching
KV Cache Management
Architecture Design
Performance Engineering
Machine Learning

Job description

Hewlett Packard Enterprise seeks a senior architect to define and own the technical architecture for LLM inference runtimes, focusing on execution efficiency and distributed inferencing strategies on customer-owned hardware. You will lead design reviews and mentor engineers.

You will collaborate with performance teams to optimize latency and throughput, apply continuous batching, and advance GPU memory management and inference internals in a fast-paced AI systems environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer (Hybrid)
Senior LLM Inference Engineer (Hybrid)

Hobbsnews • Spring (TX), Northern (KY)

Hybrid
USD 137,000 - 315,000
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python
Principal LLM Inference Architect – Kubernetes, CUDA & Go/Python

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Senior LLM Inference Engineer — Hybrid (Go/Python)
Senior LLM Inference Engineer — Hybrid (Go/Python)

Hewlett Packard Enterprise Development LP • Spring (TX), Northern (KY)

Hybrid
USD 144,000 - 273,000
LLM Inference Runtime Architect for AI Accelerator
LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
Senior AI Platform Architect — Generative LLMs
Senior AI Platform Architect — Generative LLMs

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 180,000 - 280,000
Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000