Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Together AI is building state-of-the-art LLM inference infrastructure to enable scalable, low-latency deployment for multimodal models. You will design and optimize distributed inference engines, collaborate with hardware and research teams, and push performance, scalability and cost-efficiency.
We seek an engineer with 3+ years in deep learning inference or HPC, proficient in Python and C++/CUDA, and hands-on experience with TensorRT, MoE and GPU optimization.
Together AI is building state-of-the-art LLM inference infrastructure to enable scalable, low-latency deployment for multimodal models. You will design and optimize distributed inference engines, collaborate with hardware and research teams, and push performance, scalability and cost-efficiency.
We seek an engineer with 3+ years in deep learning inference or HPC, proficient in Python and C++/CUDA, and hands-on experience with TensorRT, MoE and GPU optimization.