Stand out for this role — generate a tailored resume and cover letter in about a minute.
Together AI, the AI Native Cloud, is seeking an Inference Research Intern to dive into distributed inference, compiler-aware optimization, and novel inference-time strategies. You will co-design cross-layer optimizations across models, systems, and hardware, focusing on KV cache design and large-scale serving architectures.
The internship runs 12 to 14 weeks, from January 4th to April 9th, with housing stipends and competitive benefits, and compensation at $58 to $70 per hour depending on
Together AI, the AI Native Cloud, is seeking an Inference Research Intern to dive into distributed inference, compiler-aware optimization, and novel inference-time strategies. You will co-design cross-layer optimizations across models, systems, and hardware, focusing on KV cache design and large-scale serving architectures.
The internship runs 12 to 14 weeks, from January 4th to April 9th, with housing stipends and competitive benefits, and compensation at $58 to $70 per hour depending on