Senior AI Engineer Gateworth Group On-site Fast Track available

HireHouse

Dubai

On-site

AED 400,000 - 800,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Gateworth Group partners with a technology organisation in the UAE to scale its AI engineering function, focusing on high‑performance model optimisation and advanced AI infrastructure. The role involves deep work on transformer models, inference tuning, and deployment strategies across large-scale environments.

You will contribute to analysis, tuning, benchmarking and architectural decisions, working across hardware, systems and software to push the limits of production AI systems.

Qualifications

  • Strong understanding of transformer architectures, LLM internals and both dense/MoE models.
  • Hands-on experience with modern LLMs and attention-level optimisation.
  • Background in distributed systems and large-scale inference workloads.

Responsibilities

  • Improve and optimise LLM inference performance across distributed environments.
  • Benchmark LLMs across varied hardware stacks.
  • Design attention-level optimisation (Flash Attention, grouped-query, sliding-window).
  • Deliver model-level optimisation including quantisation, KV-cache strategies, batching and parallelism.
  • Collaborate with hardware, systems and compiler teams on inference pipelines.
  • Build and maintain benchmarking frameworks to measure latency and throughput.

Skills

Transformer architectures
LLM internals
Dense & MoE models
Python (PyTorch/JAX)
Profiling performance
Distributed systems

Job description

Position: Senior AI Engineer

Compensation: Competitive salary plus family benefits & variable

Location: Dubai, United Arab Emirates

Overview

Gateworth Group is partnering with a technology organisation in the UAE that is scaling its AI engineering function and building a specialist team focused on high‑performance model optimisation. They’re investing heavily in advanced AI infrastructure and are seeking senior engineers who can work deep inside modern LLMs, improve inference behaviour, and shape how large‑scale models run in production.

This role suits someone who enjoys complex, hands‑on engineering, understands how transformer architectures behave at scale, and can move confidently between model internals, systems performance, and deployment‑level optimisation. You’ll work across analysis, tuning, benchmarking and architectural decision‑making, contributing to next‑generation AI systems.

Main Responsibilities
  • Improve and optimise LLM inference performance across distributed, multi‑chip and multi‑node environments
  • Apply strong understanding of transformer architectures, including dense and Mixture‑of‑Experts (MoE) models
  • Benchmark leading LLMs (LLaMA, Mistral, Qwen, DeepSeek) across varied hardware stacks
  • Design and implement attention‑level optimisations (Flash Attention, grouped‑query, sliding‑window)
  • Deliver model‑level optimisation including quantisation (INT8/FP8), KV‑cache strategies, batching and parallelism
  • Work closely with hardware, systems and compiler teams to co‑design efficient inference pipelines
  • Build and maintain benchmarking frameworks to measure latency, throughput and scaling behaviour
  • Evaluate architectural trade‑offs and contribute to deployment strategies for large‑scale environments
  • Stay current with research across LLM architectures, inference optimisation and performance engineering
Qualifications
  • Strong understanding of transformer architectures, LLM internals and both dense/MoE models
  • Hands‑on experience with modern LLMs (LLaMA, Mistral, Qwen, DeepSeek) and attention‑level optimisation
  • Practical experience with inference optimisation: quantisation (INT8/FP8), KV‑cache strategies, batching, pruning and parallelism
  • Strong Python skills with PyTorch or JAX, plus experience profiling and debugging performance bottlenecks
  • Background in distributed systems, large‑scale inference workloads and system‑level optimisation across hardware and runtime layers
  • Ideally 8+ years in deep learning, AI systems or performance engineering, with exposure to datacenter‑scale inference (e.g., vLLM) and hardware‑aware optimisation
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Dubai-based Senior AI Engineer—LLM Inference & Performance
Dubai-based Senior AI Engineer—LLM Inference & Performance

HireHouse • Dubai

On-site
AED 400,000 - 800,000
Senior AI Engineer – LLM Systems
Senior AI Engineer – LLM Systems

Evollabs • Dubai

On-site
AED 661,000 - 1,028,000
Principal AI Ops Engineer
Principal AI Ops Engineer

Discovered MENA • Abu Dhabi

On-site
AED 350,000 - 650,000
AI LLM Engineer
AI LLM Engineer

DiceTek UAE • Dubai

On-site
AED 279,000 - 446,400
Exposure to advanced AI technologies
Opportunity for career growth in AI engineering
Work in a dynamic tech industry
Senior AI Engineer
Senior AI Engineer

Reqiva • Dubai

Hybrid
AED 335,000 - 446,000
Agentic AI Engineer | Systems Ltd | Dubai, UAE
Agentic AI Engineer | Systems Ltd | Dubai, UAE

Systems Ltd • Dubai

On-site
AED 360,000 - 600,000
Principal AI Ops Engineer
Principal AI Ops Engineer

Tanqeeb • Abu Dhabi

On-site
AED 320,000 - 480,000
Senior AI Engineer
Senior AI Engineer

Client of Reqiva • Dubai

On-site
AED 420,000 - 720,000
Artificial intelligence (AI) Engineer
Artificial intelligence (AI) Engineer

Baker Tilly JFC Group • United Arab Emirates

On-site
AED 183,621 - 257,069
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Client of Discovered MENA • Dubai

On-site
AED 350,000 - 520,000