Turn this role into an interview — a resume and cover letter built around what this employer wants.
Hewlett Packard Enterprise seeks a senior architect to define and own the technical architecture for LLM inference runtimes, focusing on execution efficiency and distributed inferencing strategies on customer-owned hardware. You will lead design reviews and mentor engineers.
You will collaborate with performance teams to optimize latency and throughput, apply continuous batching, and advance GPU memory management and inference internals in a fast-paced AI systems environment.
Define and own the technical architecture for LLM inference runtimes, focusing on execution efficiency and distributed inferencing strategies. Lead design reviews, mentor engineers, and partner with performance teams to optimize latency and throughput on customer-owned hardware.
Requires at least 12 years of software engineering experience with a minimum of 1 year working directly on LLM inference runtimes. Candidates must possess expert-level proficiency in Kubernetes, Python, Go, and deep knowledge of GPU memory hierarchies and inference internals.