Turn this role into an interview — a resume and cover letter built around what this employer wants.
Tether.io is seeking an experienced AI researcher to advance compression for multimodal systems, including LLMs and VLMs. You will reduce model footprint and compute cost while preserving accuracy across text, image, and audio streams.
You will lead quantization, distillation, and pruning efforts, publish findings, and collaborate across teams to push state-of-the-art in model efficiency for edge deployment.
As a member of our AI research team, you will drive innovation in model compression and efficient deployment for advanced multimodal AI systems, including large language models (LLMs) and vision-language models (VLMs). Your work will focus on reducing model footprint and computational cost while preserving accuracy, enabling high-performance AI to run efficiently across resource‑constrained edge devices. You will apply and advance compression techniques such as quantization, knowledge distillation, and pruning to streamline complex multimodal architectures that integrate text, images, and audio.