PhD Intern: HPC & LLM Inference for Distributed Systems
Wayne State University
Richland (WA)
Hybrid
USD 34,000 - 55,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A research institution is seeking a PhD intern for Spring 2026 with a focus on High Performance Computing and Inference of Large Language Models. The candidate will engage in designing efficient cache management and developing strategies for LLM inference. Preferred qualifications include experience with open-source inference engines and GPU profilers. This internship offers flexibility in duration with a minimum of 3 months and can be remote or onsite.
Qualifications
Candidate must have a strong background in High Performance Computing.
Research experience demonstrated is preferred.
Familiarity with Fabric Attached Memory protocols is a plus.
Responsibilities
Design efficient KV cache management for distributed LLM Inference.
Develop hybrid parallelism strategies for LLM inference.
Participate in the development and publication of research.
Skills
High Performance Computing
Distributed LLM inference
Open-source inference engines
GPU profilers
CXL protocols
Education
Currently enrolled in a PhD program
Minimum GPA of 3.0
Job description
A research institution is seeking a PhD intern for Spring 2026 with a focus on High Performance Computing and Inference of Large Language Models. The candidate will engage in designing efficient cache management and developing strategies for LLM inference. Preferred qualifications include experience with open-source inference engines and GPU profilers. This internship offers flexibility in duration with a minimum of 3 months and can be remote or onsite.