Role Overview
At YTL AI Labs, we build sovereign AI models that perform on par with the world’s best—while staying grounded in local needs, values, and context. Our flagship model, Ilmu, is designed to be culturally aware, contextually intelligent, and fluent in Bahasa Melayu, delivering cutting-edge solutions that empower Malaysian businesses with intelligence that truly understands the market and the people they serve.
As pioneers of sovereign AI, we believe every nation should have the power to shape its own intelligence—guided by its people, priorities, and principles.
We are seeking a Model Performance Engineer to lead the design and optimization of large-scale AI inference systems.
This is a high-impact role at the intersection of machine learning research and distributed systems engineering, where you will drive how frontier models are deployed, scaled, and experienced in production.
You will play a key role in shaping our inference architecture, pushing the limits of latency, throughput, and cost efficiency, and translating cutting-edge research into robust, real-world systems.
Key Responsibilities
- Lead the design and implementation of high-performance inference systems for LLMs, multimodal, and speech models
- Drive improvements in:
- Throughput (tokens/sec, QPS)
- Architect and optimize distributed inference systems across GPU clusters
- Own and implement advanced techniques such as:
- KV-cache optimization and memory management
- Evaluate and integrate serving frameworks (vLLM, TensorRT-LLM, Triton, custom runtimes)
- Partner with research teams to productionize new models and architectures
- Build and maintain benchmarking and evaluation pipelines for inference performance
- Diagnose and resolve bottlenecks across:
- Networking and distributed systems
- Mentor engineers and contribute to technical direction and best practices
Key Skills and Qualifications:
Core Requirements
- 3–5+ years of experience in ML systems, backend engineering, or high-performance computing
- Strong expertise in:
- Deep learning frameworks (PyTorch, JAX, TensorFlow)
- Python and/or C++
- Deep understanding of:
- Transformer architectures and modern LLM systems
- GPU architecture and parallel computing
Experience
- Proven track record building or optimizing large-scale inference systems in production
- Hands‑on experience with:
- GPU optimization (CUDA, kernel tuning, memory management)
- Experience scaling systems handling high concurrency workloads
Nice to Have
- Experience with compiler stacks (XLA, TVM, MLIR)
- Familiarity with hardware accelerators (NVIDIA, AMD, TPUs)
- Contributions to open-source ML systems
- Background in multimodal or speech model serving
- Published research in ML systems or efficiency
What Sets You Apart
- You instinctively think in:
- tokens/sec, GPU utilization, and tail latency
- You can bridge:
- You are comfortable operating across the stack:
- You take ownership of performance as a product feature
- Own critical parts of the inference stack and roadmap
- Drive technical decisions and trade-offs across teams
- Mentor junior engineers and elevate team standards
- Influence system design across model, infra, and product layers
Why Join Us
- Work on frontier AI systems deployed at scale
- Solve deeply technical challenges in efficiency and systems design
- Be part of building sovereign AI infrastructure
- Shape how AI reaches millions of users in real-world applications