Senior Software Engineer, Inference

Hewlett Packard Enterprise

Durham (NC)

On-site

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health benefits
Professional development programs
Flexible work arrangements

Job summary

Hewlett Packard Enterprise is seeking an experienced software engineer to design and implement core components of the LLM runtime, focusing on engine integration, batching, and KV cache management. You will partner with distributed teams to optimize latency, throughput, and execution across deployments.

The ideal candidate has 8+ years in software engineering with 1-2+ years in LLM inference runtimes, and strong expertise in Go, Python, Kubernetes, and deep engine internals.

Qualifications

  • Requires at least 8 years of software engineering experience with 1-2+ years in LLM inference runtimes or production model serving.
  • Strong proficiency in Go, Python, Kubernetes, and deep knowledge of inference engine internals.
  • Experience optimizing latency, throughput, and distributed execution for AI workloads.

Responsibilities

  • Design and implement core components of the LLM runtime, including engine integration and batching.
  • Collaborate with engineering teams to optimize inference performance and KV cache management.

Skills

LLM inference
Go
Python
Kubernetes
C++
CUDA
TensorRT-LLM
vLLM
Tensor parallelism
Pipeline parallelism
Performance profiling
Distributed execution

Tools

Kubernetes
Python
Go
C++
CUDA
TensorRT-LLM
vLLM
Distributed inference tooling

Job description

Design and implement key components of the LLM runtime, including engine integration, batching, and KV cache management. Partner with engineering teams to optimize inference performance, including latency, throughput, and distributed execution capabilities.

Requirements: Requires at least 8 years of software engineering experience with 1-2+ years specifically in LLM inference runtimes or production model serving. Candidates must possess strong proficiency in Go, Python, Kubernetes, and deep knowledge of inference engine internals.

Key Skills: LLM inference, Kubernetes, Go, Python, C++, CUDA, vLLM, TensorRT-LLM, Continuous batching, KV cache management, Distributed execution, GPU optimization, Tensor parallelism, Pipeline parallelism, Performance profiling

Benefits: Health and wellbeing benefits, Professional development programs, Flexible work arrangements

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Senior LLM Inference Runtime Engineer
Senior LLM Inference Runtime Engineer

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 150,000 - 210,000
Health benefits
Professional development programs
Flexible work arrangements
Senior Software Engineer - AI Inference
Senior Software Engineer - AI Inference

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
Staff Software Engineer, AI Inference
Staff Software Engineer, AI Inference

ChatGPT Jobs • New York (NY)

On-site
USD 180,000 - 240,000
Health Insurance
Equity
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis