Senior AI Engineer – LLM Systems

Evollabs Tech

Dubai

On-site

AED 420,000 - 660,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evollabs Tech in Dubai is seeking an experienced engineer to advance LLM inference performance across distributed, multi-chip platforms. You will optimize transformer-based models and collaborate with hardware teams to push efficiency in AI inference at datacenter scale.

The role focuses on state-of-the-art LLM internals, quantization, KV caching, and parallelism strategies, with hands-on work on LLaMA, Mistral, Qwen, and related technologies.

Qualifications

  • Strong understanding of transformer architectures and LLM internals.
  • Hands-on experience with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen).
  • Experience with LLM inference optimization techniques (quantization, KV caching, batching).

Responsibilities

  • Analyze, profile and optimize LLM inference performance across distributed, multi-chip systems.
  • Evaluate and benchmark different LLMs on custom hardware.
  • Design and implement optimizations for attention mechanisms and model-level performance.

Skills

Transformer architectures
LLM internals
Performance optimization
PyTorch/JAX
Distributed systems
Python
Benchmarking

Tools

PyTorch
JAX
CUDA

Job description

Description

We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.

Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand how modern large language models behave at scale.

This role is focused on LLM systems, architecture, and performance. You will work at the intersection of model internals and hardware, ensuring that state-of-the-art models run efficiently on our platform.

Responsibilities
  • Analyze, profile and optimize large language model (LLM) inference performance across distributed, multi-chip systems
  • Bring a deep understanding of transformer architectures, including dense and Mixture-of-Experts (MoE) models
  • Evaluate and benchmark different LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek) on custom hardware
  • Design and implement optimizations for attention mechanisms (e.g., Flash Attention, grouped-query attention, sliding window attention)
  • Work on model-level optimizations such as quantization (INT8/FP8), KV-cache management, batching, and parallelism strategies
  • Collaborate with hardware and compiler teams to co-design efficient inference pipelines
  • Build and maintain benchmarking frameworks to evaluate latency, throughput, and scaling behavior
  • Analyze trade-offs between model architecture choices and system-level performance
  • Contribute to model deployment strategies for large-scale datacenter environments
  • Stay up to date with the latest research in LLM architectures and inference optimization
Requirements
  • Strong understanding of transformer architectures and LLM internals
  • Hands-on experience working with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek, etc.)
  • Deep knowledge of different LLM architectures such as dense models and Mixture-of-Experts (MoE) architectures
  • Familiarity with attention mechanisms and their optimizations
  • Experience with LLM inference optimization techniques (quantization, pruning, KV caching, batching, etc.)
  • Strong Python skills and experience with ML frameworks (PyTorch, JAX, or similar)
  • Experience with distributed systems and large-scale inference workloads
  • Ability to profile and debug performance bottlenecks across hardware and software stacks
  • Strong systems thinking and ability to work across model, runtime, and hardware layers
Preferred Qualifications
  • 8+ years of relevant experience in deep learning, AI systems, or performance engineering
  • Experience working close to hardware (GPU, TPU, or custom accelerators)
  • Experience with parallelism strategies (tensor parallelism, pipeline parallelism, expert parallelism)
  • Familiarity with datacenter-scale deployment, orchestration and inference servers (e.g., vLLM)
  • Background in performance engineering or systems optimization
What We’re Not Looking For

This role is not focused on prompt engineering, or application-layer GenAI development. Instead, it is centered on deep LLM internals, architecture, and inference performance at scale.

Why Join Us?
  • Work on cutting-edge AI hardware designed specifically for LLM inference
  • Solve challenging problems at the intersection of AI models and systems
  • Collaborate with a team pushing the boundaries of datacenter-scale AI performance
  • Make a direct impact on the future of AI infrastructure
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Systems Engineer: Scale Inference
Senior LLM Systems Engineer: Scale Inference

Evollabs Tech • Dubai

On-site
AED 420,000 - 660,000
Principal AI – Large Language Models
Principal AI – Large Language Models

Uney GmbH • Dubai

On-site
AED 300,000 - 450,000
Competitive salary
Comprehensive health benefits
Collaborative work environment
AI Engineer (Applied)
AI Engineer (Applied)

BlackStone eIT • Dubai

On-site
AED 120,000 - 150,000
Paid Time Off
Performance Bonus
Training & Development
AI Engineer
AI Engineer

SUPERBOT (BOT) • Dubai

On-site
AED 180,000 - 300,000
Visa processing
UAE mandated benefits
Generous holidays
+2
Applied AI Engineer
Applied AI Engineer

IDCMPS • Abu Dhabi

On-site
Lead AI Scientist / Head of AI Solutions
Lead AI Scientist / Head of AI Solutions

Recenso • Abu Dhabi

On-site
Opportunity to lead AI innovation
Research-driven environment
Leadership exposure across teams
Staff Python / PyTorch Developer — Frontend Inference Compiler – Dubai
Staff Python / PyTorch Developer — Frontend Inference Compiler – Dubai

Cerebras Systems, Inc. • Dubai

On-site
AED 367,000 - 515,000
Equal opportunity employer
Simple, non-corporate work culture
Support for continuous learning and growth
Staff Python / PyTorch Developer — Frontend Inference Compiler - Dubai
Staff Python / PyTorch Developer — Frontend Inference Compiler - Dubai

Cerebras • United Arab Emirates

On-site
Diverse and inclusive work environment
Opportunities for continuous learning and growth
Non-corporate work culture
AI Engineer, LLM & Agentic Systems
AI Engineer, LLM & Agentic Systems

Aumnitech • Dubai

On-site
AED 240,000 - 420,000
AI LLM Engineer
AI LLM Engineer

DiceTek UAE • Dubai

On-site
Exposure to advanced AI technologies
Opportunity for career growth in AI engineering
Work in a dynamic tech industry