Get more replies from employers
Send a job-specific resume in minutes.
Together AI is building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). We seek an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale.
This role focuses on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design for scalable LLM and vision model deployment.
Together AI is building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). We seek an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale.
This role focuses on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design for scalable LLM and vision model deployment.