Get more replies from employers
Send a job-specific resume in minutes.
Evollabs Tech in Dubai is seeking an experienced engineer to advance LLM inference performance across distributed, multi-chip platforms. You will optimize transformer-based models and collaborate with hardware teams to push efficiency in AI inference at datacenter scale.
The role focuses on state-of-the-art LLM internals, quantization, KV caching, and parallelism strategies, with hands-on work on LLaMA, Mistral, Qwen, and related technologies.
Evollabs Tech in Dubai is seeking an experienced engineer to advance LLM inference performance across distributed, multi-chip platforms. You will optimize transformer-based models and collaborate with hardware teams to push efficiency in AI inference at datacenter scale.
The role focuses on state-of-the-art LLM internals, quantization, KV caching, and parallelism strategies, with hands-on work on LLaMA, Mistral, Qwen, and related technologies.