LLM Systems Engineer — High-Performance Inference

Evollabs

Dubai

On-site

AED 661,000 - 1,028,000

Full time

38 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Evollabs is building the next generation of AI infrastructure and seeks engineers focused on LLM systems, architecture, and performance. You will work at the intersection of model internals and hardware to ensure state-of-the-art models run efficiently on our platform.

Responsibilities include profiling, benchmarking, and optimizing LLM inference across distributed systems, designing attention and quantization optimizations, and collaborating with hardware teams to co-design efficient pipelines.

Qualifications

  • Strong understanding of transformer architectures and LLM internals.
  • Hands-on experience with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek).
  • Deep knowledge of MoE architectures and attention optimizations.

Responsibilities

  • Analyze, profile and optimize LLM inference performance across distributed, multi-chip systems.
  • Understand transformer architectures including dense and MoE models.
  • Evaluate and benchmark LLMs on custom hardware (e.g., LLaMA, Mistral, Qwen, DeepSeek).
  • Design optimizations for attention mechanisms (Flash Attention, grouped-query, sliding window).
  • Work on model-level optimizations: quantization (INT8/FP8), KV-cache management, batching, parallelism strategies.
  • Collaborate with hardware and compiler teams to co-design inference pipelines.
  • Build and maintain benchmarking frameworks for latency, throughput, scaling.
  • Analyze trade-offs between model architecture choices and system-level performance.
  • Stay up to date with latest LLM research and inference optimization.

Skills

Transformer architectures
LLM inference optimization
Python & ML frameworks

Tools

PyTorch
JAX

Job description

Evollabs is building the next generation of AI infrastructure and seeks engineers focused on LLM systems, architecture, and performance. You will work at the intersection of model internals and hardware to ensure state-of-the-art models run efficiently on our platform.

Responsibilities include profiling, benchmarking, and optimizing LLM inference across distributed systems, designing attention and quantization optimizations, and collaborating with hardware teams to co-design efficient pipelines.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Engineer – LLM Systems
Senior AI Engineer – LLM Systems

Evollabs • Dubai

On-site
AED 661,000 - 1,028,000
Engineering Manager, AI/LLM Systems & Delivery
Engineering Manager, AI/LLM Systems & Delivery

Hire Feed • United Arab Emirates

On-site
AED 300,000 - 540,000
AI Engineering Lead for LLM & ML Systems
AI Engineering Lead for LLM & ML Systems

AI71 • Abu Dhabi

On-site
AED 360,000 - 540,000
Flexible work environment
Competitive compensation
Career growth
+1
AI Engineer
AI Engineer

Tanqeeb • Abu Dhabi

On-site
AED 350,000 - 600,000
Dubai Onsite Agentic AI Engineer: Ultra-Low Latency LLMs
Dubai Onsite Agentic AI Engineer: Ultra-Low Latency LLMs

Systems Ltd • Dubai

On-site
AED 360,000 - 600,000
Senior Generative AI Engineer—LLM & MLOps
Senior Generative AI Engineer—LLM & MLOps

Tanqeeb • Abu Dhabi

On-site
AED 350,000 - 600,000
AI Engineer
AI Engineer

Cognitive • Abu Dhabi

On-site
AED 350,000 - 650,000
AI Engineer Internship – LLM Data
AI Engineer Internship – LLM Data

Institute of Foundation Models • Abu Dhabi

On-site
AED 17,000 - 33,000
Senior AI Production Engineer: LLMs & MLOps
Senior AI Production Engineer: LLMs & MLOps

Dicetek LLC • Dubai

On-site
AED 360,000 - 600,000
LLM Production Architect & AI Transformation Lead
LLM Production Architect & AI Transformation Lead

finera. • Dubai

On-site
AED 350,000 - 700,000
Medical Insurance starting day 1
Training resources
Well-stocked office
+2