Lead AI Inference & Performance Engineer

Dizzaract

Abu Dhabi

On-site

AED 420,000 - 720,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
24 days annual leave
Modern office in Yas Creative Hub

Job summary

Dizzaract is seeking a Lead AI Engineer for Inference Serving & Performance to own the technical direction of our inference stack. You will work across inference serving, GPU performance, distributed systems, benchmarking, and model optimisation to accelerate serving infrastructure while preserving quality.

This senior role requires hands‑on leadership, setting engineering standards, and guiding teams across serving, metrics, and testing.

Qualifications

  • Strong hands‑on experience building, optimising, or operating LLM inference‑serving systems in production.
  • Deep understanding of inference stacks such as vLLM, SGLang, TensorRT‑LLM, or similar.
  • Understanding of paged attention, continuous batching, disaggregated serving, KV-cache management, and modern inference architectures.
  • Experience with CUDA or Triton for GPU performance engineering.
  • Experience with MoE models, expert parallelism, and distributed execution.
  • Experience with model quantisation and measuring quality against full‑precision references.
  • Strong benchmarking experience including load generation and latency percentiles.
  • Distributed systems knowledge for routing, scheduling, and workload distribution.
  • Proven technical leadership in inference serving or ML infrastructure.

Responsibilities

  • Inference Serving: Own and improve the FAR Labs inference serving stack and related optimisations.
  • Serving Performance: Improve throughput, latency, memory efficiency, and cost while maintaining model quality.
  • GPU Performance Engineering: Identify bottlenecks across the stack using CUDA, Triton, profiling, and memory optimisations.
  • Mixture-of-Experts Serving: Drive efficient serving of MoE models with expert parallelism and load balancing.
  • Quantization & Model Optimisation: Implement quantization and measure against full‑precision references.
  • Benchmarking & Measurement: Own reproducible benchmarking across performance, quality, and cost.
  • Distributed Serving: Make routing, scheduling, and caching decisions across diverse hardware environments.
  • Technical Roadmap: Define serving and measurement roadmap and prioritise improvements.
  • Engineering Standards: Establish standards around inference performance, reproducibility, and production readiness.
  • Technical Leadership: Mentor engineers and influence architecture while remaining hands‑on.

Skills

Inference serving
Performance optimization
GPU utilisation
Distributed systems
Benchmarking & measurement
Technical leadership
Speculative decoding
English (Advanced)

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton

Job description

Dizzaract is seeking a Lead AI Engineer for Inference Serving & Performance to own the technical direction of our inference stack. You will work across inference serving, GPU performance, distributed systems, benchmarking, and model optimisation to accelerate serving infrastructure while preserving quality.

This senior role requires hands‑on leadership, setting engineering standards, and guiding teams across serving, metrics, and testing.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Platform PM — Remote
Senior AI Inference Platform PM — Remote

Dizzaract • United Arab Emirates

On-site
AED 350,000 - 520,000
Health insurance
24 days annual leave
Modern office in Yas Creative Hub
Engineering Leader - AI & Gaming Systems
Engineering Leader - AI & Gaming Systems

Dizzaract • Abu Dhabi

On-site
AED 420,000 - 660,000
Tech Lead - AI-Driven Gaming Platform
Tech Lead - AI-Driven Gaming Platform

Dizzaract • Abu Dhabi

On-site
AED 320,000 - 420,000
Health insurance
24 days annual leave
Modern office in Yas Creative Hub
Lead AI Engineer (Inference Serving & Performance)
Lead AI Engineer (Inference Serving & Performance)

Dizzaract • Abu Dhabi

On-site
AED 420,000 - 720,000
Health insurance
24 days annual leave
Modern office in Yas Creative Hub
LLM Systems Engineer — High-Performance Inference
LLM Systems Engineer — High-Performance Inference

Evollabs • Dubai

On-site
AED 661,000 - 1,028,000
Senior AI Engineer – LLM Systems
Senior AI Engineer – LLM Systems

Evollabs • Dubai

On-site
AED 661,000 - 1,028,000
Principal AI Ops Engineer
Principal AI Ops Engineer

Discovered MENA • Abu Dhabi

On-site
AED 350,000 - 650,000
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • United Arab Emirates

Remote
AED 350,000 - 700,000
Remote-first
International team
Cutting-edge research
+2
AI Inference Architect: Kernel & Edge Optimizations
AI Inference Architect: Kernel & Edge Optimizations

Lever, Inc. • United Arab Emirates

Remote
AED 350,000 - 700,000
Remote-first
International team
Cutting-edge research
+2
Senior Backend Engineer - Go, APIs & AI Platforms
Senior Backend Engineer - Go, APIs & AI Platforms

Dizzaract • Abu Dhabi

On-site
AED 300,000 - 420,000
Health insurance
24 days annual leave
Modern office in Yas Creative Hub