Location: Hong Kong (Onsite)
Company: Xircuit AI
About Xircuit AI
Xircuit is building infrastructure to automatically optimize LLM inference performance across different models, workloads, and hardware.
We work across the inference stack—from GPU kernels and quantization to compilation, batching, scheduling, and serving engines—to improve latency, throughput, memory usage, and cost per token.
The Role
We are looking for an LLM Inference Optimization Engineer to join our early technical team.
You will profile and optimize real LLM workloads, develop new optimization techniques,design experiments and benchmarks inspired by research papers, and help build AI agents that automate parts of the performance-engineering and experimentation workflow.
Core Skill Requirements
- Python: Strong Python skills, including PyTorch, numerical computing, profiling, testing, and performance-oriented code.
- LLM Inference: Experience with vLLM, SGLang, TensorRT-LLM, PyTorch, or similar systems.
- GPU Performance Engineering: Understanding of CUDA, GPU memory, parallelism, profiling, and kernel performance.
- Kernel Optimization: Experience with CUDA, Triton, kernel fusion, autotuning, or compiler-generated kernels.
- AI Agents: Experience building tool-using LLM agents and workflows with frameworks such as LangGraph, LangChain, OpenAI Agents SDK, or similar orchestration frameworks.
- Model Efficiency: Knowledge of quantization and low-precision inference such as, BF16, FP8 and FP4.
- Performance Analysis: Ability to design rigorous experiments and validate whether optimizations produce real end-to-end improvements.
Bonus Skills
Experience in any of the following is a strong plus:
- Cross-Vendor Hardware: Optimizing workloads beyond NVIDIA, including AMD GPUs/ROCm, NPUs, TPUs, and other AI accelerators.
- Portable Kernel Development: Experience with hardware-portable approaches such as Triton, HIP, MLIR, OpenXLA, TVM, or similar compiler stacks.
- Inference Runtime Internals: CUDA Graphs, TorchInductor, KV-cache optimization, speculative decoding, batching, and scheduling.
- Automated Optimization: Search, autotuning, reinforcement learning, evolutionary methods, or agent-based systems for discovering performance improvements.
- Relevant Research Publications: Publications in LLM inference, GPU/kernel optimization, compilers, systems for ML, quantization, or heterogeneous AI hardware are a strong plus.
Soft Skills & Attitude
- Experimental Rigor: You care about reproducible benchmarks and measurable improvements.
- First-Principles Thinking: You enjoy understanding why systems are slow and finding unconventional solutions.
- Research Mindset: You are comfortable exploring new ideas and learning from failed experiments.
- Founder’s Mindset: You are proactive, independent, and comfortable in an early-stage environment with a strong sense of commitment.
What We Offer
- Competitive salary
- Founding-team impact on Xircuit's technical direction
- Hybrid flexibility in Hong Kong
- Access to diverse AI hardware
- Work across LLM inference, GPU/NPU/TPU optimization, AI agents, compilers, quantization, and automated performance engineering