Model Performance Engineer

YTL AI Labs

Kuala Lumpur

On-site

MYR 180,000 - 300,000

Full time

42 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

YTL AI Labs, a Malaysian sovereign AI company, seeks a Model Performance Engineer to design and optimize large-scale inference systems for LLMs, multimodal, and speech models. You will drive latency, throughput, and cost efficiency while productionizing cutting-edge research.

You will work across GPU clusters, memory management, and serving frameworks to ensure robust, real-world deployment at scale.

Qualifications

  • 3–5+ years in ML systems, backend or HPC.
  • Strong experience with large-scale inference systems in production.
  • Hands-on GPU optimization and memory management.
  • Familiar with serving frameworks (vLLM, TensorRT-LLM, Triton).

Responsibilities

  • Lead design and implementation of high-performance inference systems for LLMs and multimodal models.
  • Optimize throughput, latency, and memory usage across GPU clusters.
  • Collaborate with research to productionize new models and architectures.
  • Build benchmarking pipelines and diagnose bottlenecks in distributed systems.

Skills

Deep learning frameworks
Python
C++
GPU architectures
Distributed systems

Tools

CUDA
TensorRT
Triton
vLLM

Job description

Role Overview

At YTL AI Labs, we build sovereign AI models that perform on par with the world’s best—while staying grounded in local needs, values, and context. Our flagship model, Ilmu, is designed to be culturally aware, contextually intelligent, and fluent in Bahasa Melayu, delivering cutting-edge solutions that empower Malaysian businesses with intelligence that truly understands the market and the people they serve.

As pioneers of sovereign AI, we believe every nation should have the power to shape its own intelligence—guided by its people, priorities, and principles.

We are seeking a Model Performance Engineer to lead the design and optimization of large-scale AI inference systems.

This is a high-impact role at the intersection of machine learning research and distributed systems engineering, where you will drive how frontier models are deployed, scaled, and experienced in production.

You will play a key role in shaping our inference architecture, pushing the limits of latency, throughput, and cost efficiency, and translating cutting-edge research into robust, real-world systems.

Key Responsibilities
  • Lead the design and implementation of high-performance inference systems for LLMs, multimodal, and speech models
  • Drive improvements in:
  • Throughput (tokens/sec, QPS)
  • Architect and optimize distributed inference systems across GPU clusters
  • Own and implement advanced techniques such as:
  • KV-cache optimization and memory management
  • Evaluate and integrate serving frameworks (vLLM, TensorRT-LLM, Triton, custom runtimes)
  • Partner with research teams to productionize new models and architectures
  • Build and maintain benchmarking and evaluation pipelines for inference performance
  • Diagnose and resolve bottlenecks across:
  • Networking and distributed systems
  • Mentor engineers and contribute to technical direction and best practices
Key Skills and Qualifications:
Core Requirements
  • 3–5+ years of experience in ML systems, backend engineering, or high-performance computing
  • Strong expertise in:
  • Deep learning frameworks (PyTorch, JAX, TensorFlow)
  • Python and/or C++
  • Deep understanding of:
  • Transformer architectures and modern LLM systems
  • GPU architecture and parallel computing
Experience
  • Proven track record building or optimizing large-scale inference systems in production
  • Hands‑on experience with:
  • GPU optimization (CUDA, kernel tuning, memory management)
  • Experience scaling systems handling high concurrency workloads
Nice to Have
  • Experience with compiler stacks (XLA, TVM, MLIR)
  • Familiarity with hardware accelerators (NVIDIA, AMD, TPUs)
  • Contributions to open-source ML systems
  • Background in multimodal or speech model serving
  • Published research in ML systems or efficiency
What Sets You Apart
  • You instinctively think in:
  • tokens/sec, GPU utilization, and tail latency
  • You can bridge:
  • You are comfortable operating across the stack:
  • You take ownership of performance as a product feature
  • Own critical parts of the inference stack and roadmap
  • Drive technical decisions and trade-offs across teams
  • Mentor junior engineers and elevate team standards
  • Influence system design across model, infra, and product layers
Why Join Us
  • Work on frontier AI systems deployed at scale
  • Solve deeply technical challenges in efficiency and systems design
  • Be part of building sovereign AI infrastructure
  • Shape how AI reaches millions of users in real-world applications
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Model Performance Engineer – Frontier AI Inference
Senior Model Performance Engineer – Frontier AI Inference

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior AI Engineer
Senior AI Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Senior AI Site Reliability Engineer
Senior AI Site Reliability Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 360,000
AI Site Reliability Engineer
AI Site Reliability Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 120,000 - 200,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

iSoftStone • Kuala Lumpur

On-site
MYR 60,000 - 100,000
AI Product Engineer
AI Product Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Full Stack AI Engineer
Full Stack AI Engineer

RSGx • Kuala Lumpur

Hybrid
MYR 180,000 - 240,000
Hybrid work
KL office nearby
AI Engineer
AI Engineer

Mercedes Benz • Puchong

Hybrid
MYR 120,000 - 180,000
Full Stack AI Engineer
Full Stack AI Engineer

Resource Services Group X Pty Ltd • Kuala Lumpur

Hybrid
MYR 180,000 - 320,000
Hybrid working arrangements in Kuala L
IEGG - AI Infra Engineer (3rd Party Contract -1 Year Renewable)
IEGG - AI Infra Engineer (3rd Party Contract -1 Year Renewable)

Tencent • Kuala Lumpur

On-site
MYR 120,000 - 180,000