Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Together AI in San Francisco invites a research intern to join the Inference Research team, focusing on distributed inference, compiler-aware optimization, and inference-time strategies for efficient foundation-model serving.
You will co-design cross-layer optimizations, run rigorous experiments, and publish findings while gaining exposure to CUDA, PyTorch, and large-scale ML systems within an AI-native cloud company.
Together AI in San Francisco invites a research intern to join the Inference Research team, focusing on distributed inference, compiler-aware optimization, and inference-time strategies for efficient foundation-model serving.
You will co-design cross-layer optimizations, run rigorous experiments, and publish findings while gaining exposure to CUDA, PyTorch, and large-scale ML systems within an AI-native cloud company.