PhD Intern: HPC & LLM Inference for Distributed Systems

Wayne State University

Richland (WA)

Hybrid

USD 34,000 - 55,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A research institution is seeking a PhD intern for Spring 2026 with a focus on High Performance Computing and Inference of Large Language Models. The candidate will engage in designing efficient cache management and developing strategies for LLM inference. Preferred qualifications include experience with open-source inference engines and GPU profilers. This internship offers flexibility in duration with a minimum of 3 months and can be remote or onsite.

Qualifications

  • Candidate must have a strong background in High Performance Computing.
  • Research experience demonstrated is preferred.
  • Familiarity with Fabric Attached Memory protocols is a plus.

Responsibilities

  • Design efficient KV cache management for distributed LLM Inference.
  • Develop hybrid parallelism strategies for LLM inference.
  • Participate in the development and publication of research.

Skills

High Performance Computing
Distributed LLM inference
Open-source inference engines
GPU profilers
CXL protocols

Education

Currently enrolled in a PhD program
Minimum GPA of 3.0

Job description

A research institution is seeking a PhD intern for Spring 2026 with a focus on High Performance Computing and Inference of Large Language Models. The candidate will engage in designing efficient cache management and developing strategies for LLM inference. Preferred qualifications include experience with open-source inference engines and GPU profilers. This internship offers flexibility in duration with a minimum of 3 months and can be remote or onsite.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PhD Intern - Continuum Computing at Pacific Northwest National Laboratory
PhD Intern - Continuum Computing at Pacific Northwest National Laboratory

Wayne State University • Richland (WA)

Hybrid
USD 34,000 - 55,000
PhD ML Intern - LLMs & Generative AI (Remote)
PhD ML Intern - LLMs & Generative AI (Remote)

Truveta • Seattle (WA)

Remote
Competitive compensation
Company-issued laptop and equipment
Opportunities for future full-time positions
PhD LLM Research Intern — Build & Publish Innovations
PhD LLM Research Intern — Build & Publish Innovations

Socket.dev • Santa Clara (CA)

On-site
USD 52,000 - 129,000
Intern benefits
PhD LLM Research Internship - Innovate & Collaborate
PhD LLM Research Internship - Innovate & Collaborate

NVIDIA • Santa Clara (CA)

On-site
USD 79,000 - 196,000
Intern benefits
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
PhD Research Internship — Large Language Models & Multimodal AI
PhD Research Internship — Large Language Models & Multimodal AI

Nvidia Corporation • Santa Clara (CA)

On-site
USD 52,000 - 129,000
Intern benefits
Inference Runtime Engineer for LLMs & Diffusion
Inference Runtime Engineer for LLMs & Diffusion

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
LLM Training & Inference Scientist (GPU-Optimized)
LLM Training & Inference Scientist (GPU-Optimized)

ByteDance • San Jose (CA)

On-site
USD 212,000 - 450,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3
Summer LLM Inference & Integration Engineer Intern
Summer LLM Inference & Integration Engineer Intern

Wayne State University • Salt Lake City (UT)

On-site
USD 22,000 - 28,000
Competitive monthly stipend
Networking opportunities
Training workshops
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000