Lead AI Engineer (Inference Serving & Performance)

Dizzaract

Abu Dhabi

On-site

AED 420,000 - 720,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
24 days annual leave
Modern office in Yas Creative Hub

Job summary

Dizzaract is seeking a Lead AI Engineer for Inference Serving & Performance to own the technical direction of our inference stack. You will work across inference serving, GPU performance, distributed systems, benchmarking, and model optimisation to accelerate serving infrastructure while preserving quality.

This senior role requires hands‑on leadership, setting engineering standards, and guiding teams across serving, metrics, and testing.

Qualifications

  • Strong hands‑on experience building, optimising, or operating LLM inference‑serving systems in production.
  • Deep understanding of inference stacks such as vLLM, SGLang, TensorRT‑LLM, or similar.
  • Understanding of paged attention, continuous batching, disaggregated serving, KV-cache management, and modern inference architectures.
  • Experience with CUDA or Triton for GPU performance engineering.
  • Experience with MoE models, expert parallelism, and distributed execution.
  • Experience with model quantisation and measuring quality against full‑precision references.
  • Strong benchmarking experience including load generation and latency percentiles.
  • Distributed systems knowledge for routing, scheduling, and workload distribution.
  • Proven technical leadership in inference serving or ML infrastructure.

Responsibilities

  • Inference Serving: Own and improve the FAR Labs inference serving stack and related optimisations.
  • Serving Performance: Improve throughput, latency, memory efficiency, and cost while maintaining model quality.
  • GPU Performance Engineering: Identify bottlenecks across the stack using CUDA, Triton, profiling, and memory optimisations.
  • Mixture-of-Experts Serving: Drive efficient serving of MoE models with expert parallelism and load balancing.
  • Quantization & Model Optimisation: Implement quantization and measure against full‑precision references.
  • Benchmarking & Measurement: Own reproducible benchmarking across performance, quality, and cost.
  • Distributed Serving: Make routing, scheduling, and caching decisions across diverse hardware environments.
  • Technical Roadmap: Define serving and measurement roadmap and prioritise improvements.
  • Engineering Standards: Establish standards around inference performance, reproducibility, and production readiness.
  • Technical Leadership: Mentor engineers and influence architecture while remaining hands‑on.

Skills

Inference serving
Performance optimization
GPU utilisation
Distributed systems
Benchmarking & measurement
Technical leadership
Speculative decoding
English (Advanced)

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton

Job description

Lead AI Engineer (Inference Serving & Performance)

Full time ·

01. ABOUT THE COMPANY

Dizzaract is a product-driven company operating at the intersection of gaming, digital platforms, and AI. We build and scale multiple products, including FAR Labs & Gamed — each exploring a different space, yet united by a shared approach: moving fast, staying curious, and focusing on things that people actually use.

We operate as a collaborative, non-hierarchical team where ideas are valued based on their impact, not their origin, and where AI is embedded across everything we build, from infrastructure to product decisions.

02. ABOUT THE ROLE

FAR Labs is building a distributed inference platform designed to serve large language models efficiently across diverse hardware. Serving quality — latency, throughput, memory efficiency, cost, and model quality — sits at the core of the product.

We’re looking for a Lead AI Engineer — Inference Serving & Performance to own the technical direction of our inference stack and how we measure its performance.

You will work across inference serving, GPU performance, distributed systems, benchmarking, and model optimisation to make our serving infrastructure faster and more efficient while maintaining a clear quality bar.

This is the senior technical seat for AI serving at FAR Labs. You will set the serving and measurement roadmap, establish engineering standards, and guide engineers working across serving, metrics, and testing.

The role is highly hands‑on. You’ll be expected to work directly with the inference stack, identify performance bottlenecks, implement improvements, and demonstrate their impact through rigorous and reproducible measurement.

03. WHO YOU ARE
  • Inference Expert: You have deep hands-on experience with LLM inference serving and understand what happens between a model being loaded and a production-grade token being served.
  • Performance Engineer: You think naturally in terms of latency, throughput, memory bandwidth, GPU utilisation, model quality, and cost — and understand how changes across the stack affect each of them.
  • GPU‑Native: You are comfortable working close to the hardware with CUDA or Triton and understand modern GPU architectures, memory behaviour, and precisions such as FP8 and FP4.
  • Distributed Systems Thinker: You understand the challenges of routing, scheduling, communication, caching, and workload distribution across heterogeneous compute infrastructure.
  • Measurement‑Driven: You care about proving performance improvements properly. You understand load generation, latency percentiles, throughput‑at‑SLO, reproducibility, and the importance of measuring speed alongside quality.
  • Technical Leader: You have previously led an inference, serving, or similarly complex technical initiative and can set direction that strong engineers can execute against.
  • Builder: You enjoy solving technically difficult problems in environments where the architecture, tooling, benchmarks, and engineering standards are still evolving.
04. RESPONSIBILITIES
  • Inference Serving: Own and improve the FAR Labs inference serving stack, including prefill/decode disaggregation, continuous batching, KV-cache management, cache‑aware routing, speculative decoding, and other serving optimisations.
  • Serving Performance: Improve throughput, latency, memory efficiency, and serving cost while maintaining clearly defined model‑quality standards.
  • GPU Performance Engineering: Identify and eliminate GPU performance bottlenecks across the inference stack using CUDA, Triton, profiling, kernel optimisation, memory optimisation, and appropriate precision strategies.
  • Mixture‑of‑Experts Serving: Drive efficient serving of large MoE models, including expert parallelism, all‑to‑all communication, expert‑load balancing, and distributed execution.
  • Quantization & Model Optimisation: Implement and evaluate quantization and other model optimisation techniques while measuring accuracy against appropriate full‑precision references.
  • Benchmarking & Measurement: Own rigorous, reproducible benchmarking across serving performance, quality, and cost. Establish measurement methodologies that can withstand external technical scrutiny.
  • Performance Instrumentation: Build and improve the instrumentation required to understand system behaviour, identify bottlenecks, compare configurations, and validate improvements.
  • Distributed Serving: Make technical decisions around routing, scheduling, caching, and workload distribution across diverse hardware environments.
  • Technical Roadmap: Define the AI serving and measurement roadmap and determine which performance problems and infrastructure improvements should be prioritised.
  • Engineering Standards: Establish clear technical standards around inference performance, benchmarking, reproducibility, testing, and production readiness.
  • Technical Leadership: Guide and mentor engineers working across serving, metrics, and testing while remaining deeply involved in implementation and technical problem solving.
05. REQUIREMENTS
  • Strong hands‑on experience building, optimising, or operating LLM inference‑serving systems in production.
  • Deep understanding of inference stacks such as vLLM, SGLang, TensorRT‑LLM, or similar technologies.
  • Strong understanding of paged attention, continuous batching, disaggregated serving, KV‑cache management, and modern inference architectures.
  • Strong GPU performance engineering experience using CUDA or Triton.
  • Strong understanding of GPU performance characteristics including memory bandwidth, compute utilisation, model bandwidth utilisation, FLOPs utilisation, and modern GPU precisions such as FP8 and FP4.
  • Experience serving large Mixture‑of‑Experts (MoE) models, including expert parallelism, all‑to‑all communication, and expert‑load balancing.
  • Hands‑on experience with model quantisation and measuring quality against full‑precision references.
  • Strong benchmarking experience, including load generation, latency percentiles, throughput‑at‑SLO, quality measurement, and reproducibility.
  • Strong understanding of distributed systems, particularly routing and scheduling workloads across varied hardware.
  • Ability to analyse complex performance bottlenecks across models, serving infrastructure, networking, and hardware.
  • Demonstrated technical leadership experience leading an inference, serving, ML infrastructure, or similarly complex engineering initiative.
  • Ability to set technical direction while remaining hands‑on with implementation and performance optimisation.
  • Experience with speculative decoding and understanding its behaviour under production load is a strong plus.
  • Open‑source contributions to inference or serving frameworks are a strong plus.
  • Experience serving reasoning models, long‑context models, or agentic multi‑turn workloads is a plus.
  • Experience with energy‑aware serving, including throughput‑per‑watt or power‑constrained environments, is a plus.
  • Experience with confidential computing or data‑residency‑aware serving is a plus.
  • Production experience with model compression techniques such as structured sparsity, activation sparsity, or KV compression is a plus.
  • English level: Advanced (C1).
06. WHAT YOU WILL LEARN
  • Inference at Scale: How to build and optimise production inference systems where latency, throughput, memory, model quality, hardware utilisation, and cost all need to be solved together.
  • Distributed AI Infrastructure: Deep exposure to distributed inference across diverse hardware environments and the architectural challenges involved in scheduling, routing, caching, and serving large models efficiently.
  • Frontier Serving Techniques: Hands‑on work with emerging approaches across disaggregated serving, speculative decoding, quantisation, MoE inference, model compression, and GPU optimisation.
  • Performance Engineering: How to build benchmarking and measurement systems that translate complex technical improvements into clear, reproducible performance and cost results.
  • Technical Leadership: The opportunity to define the serving architecture, engineering standards, and technical roadmap of a core FAR Labs product area.
07. WHAT WE OFFER
  • Real technical ownership of inference serving and performance across FAR Labs.
  • The opportunity to define the technical direction of a core part of the FAR Labs platform.
  • Hands‑on work with technically challenging problems across LLM inference, GPU optimisation, distributed systems, and AI infrastructure.
  • The opportunity to work with large models and emerging inference techniques in a production environment.
  • A small, senior engineering environment where technical decisions translate directly into product outcomes.
  • Fast execution, low bureaucracy, and a highly collaborative, idea‑driven team.
  • Competitive salary with performance-based incentives.
  • 24 days annual leave, plus public holidays.
  • Health insurance.
  • Modern office in Yas Creative Hub.
  • Continuous learning through real-world problem solving — not just theory.
  • The opportunity to influence both technical architecture and product direction.
  • A diverse, open‑minded team where ideas are genuinely heard.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Backend Engineer
Backend Engineer

Dizzaract • Abu Dhabi

On-site
AED 300,000 - 420,000
Health insurance
24 days annual leave
Modern office in Yas Creative Hub
Senior Engineering Manager
Senior Engineering Manager

Dizzaract • Abu Dhabi

On-site
AED 420,000 - 660,000
Senior Engineering Manager at Dizzaract FZ LLC
Senior Engineering Manager at Dizzaract FZ LLC

Dizzaract FZ LLC • Abu Dhabi

On-site
AED 279,000 - 335,000
24 days annual leave
Health insurance
Modern office in Yas Creative Hub
SRE Lead
SRE Lead

Dizzaract • Abu Dhabi

On-site
AED 420,000 - 680,000
Health insurance
24 days annual leave
Office in Yas Creative Hub
+1
Tech Lead
Tech Lead

Dizzaract • Abu Dhabi

On-site
AED 320,000 - 420,000
Health insurance
24 days annual leave
Modern office in Yas Creative Hub
Staff Machine Learning Engineer
Staff Machine Learning Engineer

GCS • Abu Dhabi

Hybrid
AED 250,000 - 420,000
ML/AI Engineer
ML/AI Engineer

SFORS • Dubai

On-site
AED 350,000 - 700,000
Housing support
Health insurance
Performance bonus
+1
MLOps Engineer
MLOps Engineer

ai71 • Abu Dhabi

On-site
AED 420,000 - 720,000
Competitive compensation
Flexible working environment
Health insurance
MLOps Engineer New Abu Dhabi, UAE
MLOps Engineer New Abu Dhabi, UAE

Greenhouse Software, Inc. • Abu Dhabi

On-site
AED 300,000 - 420,000
Associate ML Ops Engineer
Associate ML Ops Engineer

AppliedAI • Abu Dhabi

On-site
AED 156,000 - 234,000
Health insurance
Visa sponsorship
On-site Abu Dhabi HQ