Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Hewlett Packard Enterprise seeks a senior architect to define and own the technical architecture for LLM inference runtimes, focusing on execution efficiency and distributed inferencing strategies on customer-owned hardware. You will lead design reviews and mentor engineers.
You will collaborate with performance teams to optimize latency and throughput, apply continuous batching, and advance GPU memory management and inference internals in a fast-paced AI systems environment.
Hewlett Packard Enterprise seeks a senior architect to define and own the technical architecture for LLM inference runtimes, focusing on execution efficiency and distributed inferencing strategies on customer-owned hardware. You will lead design reviews and mentor engineers.
You will collaborate with performance teams to optimize latency and throughput, apply continuous batching, and advance GPU memory management and inference internals in a fast-paced AI systems environment.