Senior AI Engineer – LLM Systems

Evollabs

Dubai

On-site

AED 661,000 - 1,028,000

Full time

26 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Evollabs is building the next generation of AI infrastructure and seeks engineers focused on LLM systems, architecture, and performance. You will work at the intersection of model internals and hardware to ensure state-of-the-art models run efficiently on our platform.

Responsibilities include profiling, benchmarking, and optimizing LLM inference across distributed systems, designing attention and quantization optimizations, and collaborating with hardware teams to co-design efficient pipelines.

Qualifications

  • Strong understanding of transformer architectures and LLM internals.
  • Hands-on experience with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek).
  • Deep knowledge of MoE architectures and attention optimizations.

Responsibilities

  • Analyze, profile and optimize LLM inference performance across distributed, multi-chip systems.
  • Understand transformer architectures including dense and MoE models.
  • Evaluate and benchmark LLMs on custom hardware (e.g., LLaMA, Mistral, Qwen, DeepSeek).
  • Design optimizations for attention mechanisms (Flash Attention, grouped-query, sliding window).
  • Work on model-level optimizations: quantization (INT8/FP8), KV-cache management, batching, parallelism strategies.
  • Collaborate with hardware and compiler teams to co-design inference pipelines.
  • Build and maintain benchmarking frameworks for latency, throughput, scaling.
  • Analyze trade-offs between model architecture choices and system-level performance.
  • Stay up to date with latest LLM research and inference optimization.

Skills

Transformer architectures
LLM inference optimization
Python & ML frameworks

Tools

PyTorch
JAX

Job description

We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.

Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand how modern large language models behave at scale.

This role is focused on LLM systems, architecture, and performance. You will work at the intersection of model internals and hardware, ensuring that state-of-the-art models run efficiently on our platform.

Responsibilities
  • Analyze, profile and optimize large language model (LLM) inference performance across distributed, multi-chip systems
  • Bring a deep understanding of transformer architectures, including dense and Mixture-of-Experts (MoE) models
  • Evaluate and benchmark different LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek) on custom hardware
  • Design and implement optimizations for attention mechanisms (e.g., Flash Attention, grouped-query attention, sliding window attention)
  • Work on model-level optimizations such as quantization (INT8/FP8), KV-cache management, batching, and parallelism strategies
  • Collaborate with hardware and compiler teams to co-design efficient inference pipelines
  • Build and maintain benchmarking frameworks to evaluate latency, throughput, and scaling behavior
  • Analyze trade-offs between model architecture choices and system-level performance
  • Stay up to date with the latest research in LLM architectures and inference optimization
Requirements
  • Strong understanding of transformer architectures and LLM internals
  • Hands-on experience working with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek, etc.)
  • Deep knowledge of different LLM architectures such as dense models and Mixture-of-Experts (MoE) architectures
  • Familiarity with attention mechanisms and their optimizations
  • Experience with LLM inference optimization techniques (quantization, pruning, KV caching, batching, etc.)
  • Strong Python skills and experience with ML frameworks (PyTorch, JAX, or similar)
  • Experience with distributed systems and large-scale inference workloads
  • Ability to profile and debug performance bottlenecks across hardware and software stacks
  • Strong systems thinking and ability to work across model, runtime, and hardware layers
Preferred Qualifications
  • 8+ years of relevant experience in deep learning, AI systems, or performance engineering
  • Experience working close to hardware (GPU, TPU, or custom accelerators)
  • Experience with parallelism strategies (tensor parallelism, pipeline parallelism, expert parallelism)
  • Familiarity with datacenter-scale deployment, orchestration and inference servers (e.g., vLLM)
  • Background in performance engineering or systems optimization
What We're Not Looking For
  • This role is not focused on prompt engineering, or application-layer GenAI development. Instead, it is centered on deep LLM internals, architecture, and inference performance at scale
Why Join Us?
  • Work on cutting-edge AI hardware designed specifically for LLM inference
  • Solve challenging problems at the intersection of AI models and systems
  • Collaborate with a team pushing the boundaries of datacenter-scale AI performance
  • Make a direct impact on the future of AI infrastructure
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Systems Engineer — High-Performance Inference
LLM Systems Engineer — High-Performance Inference

Evollabs • Dubai

On-site
AED 661,000 - 1,028,000
AI Engineer
AI Engineer

Tanqeeb • Abu Dhabi

On-site
AED 350,000 - 600,000
Agentic AI Engineer | Systems Ltd | Dubai, UAE
Agentic AI Engineer | Systems Ltd | Dubai, UAE

Systems Ltd • Dubai

On-site
AED 360,000 - 600,000
AI Engineer
AI Engineer

Cognitive • Abu Dhabi

On-site
AED 350,000 - 650,000
Principal AI – Large Language Models
Principal AI – Large Language Models

Uney GmbH • Dubai

On-site
AED 300,000 - 450,000
Competitive salary
Comprehensive health benefits
Collaborative work environment
AI Engineer (Applied)
AI Engineer (Applied)

BlackStone eIT • Dubai

On-site
AED 120,000 - 150,000
Paid Time Off
Performance Bonus
Training & Development
Lead AI Scientist / Head of AI Solutions
Lead AI Scientist / Head of AI Solutions

Recenso • Abu Dhabi

On-site
AED 446,400 - 669,600
Opportunity to lead AI innovation
Research-driven environment
Leadership exposure across teams
Applied AI Engineer
Applied AI Engineer

IDCMPS • Abu Dhabi

On-site
AED 334,800 - 502,200
Specialist Artificial Intelligence
Specialist Artificial Intelligence

General Civil Aviation Authority • Abu Dhabi

On-site
AED 220,000 - 320,000
NLP Engineer
NLP Engineer

NorthBay Solutions LLC • Abu Dhabi

On-site
AED 350,000 - 700,000