LLM Inference Optimization Engineer

Xircuit AI Limited

Hong Kong

On-site

HKD 900,000 - 1,100,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid flexibility in HK
Access to diverse AI hardware
Founding-team influence
Competitive salary

Job summary

Xircuit AI is seeking an LLM Inference Optimization Engineer in Hong Kong to profile and optimize real-world LLM workloads, develop new techniques, and design rigorous experiments. You will help build AI agents that automate performance engineering workflows within an early technical team.

The role requires strong Python, PyTorch, CUDA and kernel optimization skills, plus experience with LLM inference systems. Hybrid flexibility in Hong Kong is offered.

Qualifications

  • Strong Python skills including PyTorch, profiling and performance-oriented code
  • Experience with LLM inference systems (vLLM, TensorRT-LLM, PyTorch)
  • Understanding CUDA, GPU memory, parallelism and kernel profiling
  • Experience with CUDA, Triton, kernel fusion or compiler-generated kernels
  • Experience building tool-using LLM agents and workflows (LangGraph, LangChain, OpenAI SDK)
  • Knowledge of quantization and low-precision inference (BF16, FP8, FP4)
  • Ability to design rigorous experiments and validate end-to-end gains

Responsibilities

  • Profile and optimize real LLM workloads
  • Develop new optimization techniques and experiments
  • Design benchmarks inspired by research papers
  • Help build AI agents that automate performance engineering workflows

Skills

Python & PyTorch
LLM Inference
GPU Performance
Kernel Optimization
AI Agents Frameworks
Model Efficiency (quantization)
Performance Analysis

Tools

PyTorch
CUDA
Triton
LangGraph / LangChain
OpenAI Agents SDK

Job description

Location: Hong Kong (Onsite)

Company: Xircuit AI

About Xircuit AI

Xircuit is building infrastructure to automatically optimize LLM inference performance across different models, workloads, and hardware.

We work across the inference stack—from GPU kernels and quantization to compilation, batching, scheduling, and serving engines—to improve latency, throughput, memory usage, and cost per token.

The Role

We are looking for an LLM Inference Optimization Engineer to join our early technical team.

You will profile and optimize real LLM workloads, develop new optimization techniques,design experiments and benchmarks inspired by research papers, and help build AI agents that automate parts of the performance-engineering and experimentation workflow.

Core Skill Requirements
  • Python: Strong Python skills, including PyTorch, numerical computing, profiling, testing, and performance-oriented code.
  • LLM Inference: Experience with vLLM, SGLang, TensorRT-LLM, PyTorch, or similar systems.
  • GPU Performance Engineering: Understanding of CUDA, GPU memory, parallelism, profiling, and kernel performance.
  • Kernel Optimization: Experience with CUDA, Triton, kernel fusion, autotuning, or compiler-generated kernels.
  • AI Agents: Experience building tool-using LLM agents and workflows with frameworks such as LangGraph, LangChain, OpenAI Agents SDK, or similar orchestration frameworks.
  • Model Efficiency: Knowledge of quantization and low-precision inference such as, BF16, FP8 and FP4.
  • Performance Analysis: Ability to design rigorous experiments and validate whether optimizations produce real end-to-end improvements.
Bonus Skills

Experience in any of the following is a strong plus:

  • Cross-Vendor Hardware: Optimizing workloads beyond NVIDIA, including AMD GPUs/ROCm, NPUs, TPUs, and other AI accelerators.
  • Portable Kernel Development: Experience with hardware-portable approaches such as Triton, HIP, MLIR, OpenXLA, TVM, or similar compiler stacks.
  • Inference Runtime Internals: CUDA Graphs, TorchInductor, KV-cache optimization, speculative decoding, batching, and scheduling.
  • Automated Optimization: Search, autotuning, reinforcement learning, evolutionary methods, or agent-based systems for discovering performance improvements.
  • Relevant Research Publications: Publications in LLM inference, GPU/kernel optimization, compilers, systems for ML, quantization, or heterogeneous AI hardware are a strong plus.
Soft Skills & Attitude
  • Experimental Rigor: You care about reproducible benchmarks and measurable improvements.
  • First-Principles Thinking: You enjoy understanding why systems are slow and finding unconventional solutions.
  • Research Mindset: You are comfortable exploring new ideas and learning from failed experiments.
  • Founder’s Mindset: You are proactive, independent, and comfortable in an early-stage environment with a strong sense of commitment.
What We Offer
  • Competitive salary
  • Founding-team impact on Xircuit's technical direction
  • Hybrid flexibility in Hong Kong
  • Access to diverse AI hardware
  • Work across LLM inference, GPU/NPU/TPU optimization, AI agents, compilers, quantization, and automated performance engineering
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hong Kong Onsite: LLM Inference Performance Engineer
Hong Kong Onsite: LLM Inference Performance Engineer

Xircuit AI Limited • Hong Kong

On-site
HKD 900,000 - 1,100,000
Hybrid flexibility in HK
Access to diverse AI hardware
Founding-team influence
+1
AI Engineer (LLM/ Chatbot)
AI Engineer (LLM/ Chatbot)

Pantheon Lab Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
AI Engineer / CTO
AI Engineer / CTO

Page Executive • Hong Kong

On-site
Confidential
Competitive salary package
Performance-based bonus
Opportunities to build AI from 0 - 1
+2
AI Engineer / CTO
AI Engineer / CTO

Michael Page International (Hong Kong) Limited • Hong Kong

On-site
HKD 800,000 - 1,200,000
Competitive salary package
Performance-based bonus
Opportunities to build the AI engine from 0 - 1
+1
LLM Engineer (Data and Optimization)
LLM Engineer (Data and Optimization)

TCL Corporate Research(HK) Co., Ltd • Hong Kong

Hybrid
HKD 900,000 - 1,300,000
Quant Research Engineer
Quant Research Engineer

Millennium • Hong Kong

On-site
HKD 900,000 - 1,300,000
PYTHON / AI PLATFORM ENGINEER – QUANTITATIVE TRADING & INFRASTRUCTURE
PYTHON / AI PLATFORM ENGINEER – QUANTITATIVE TRADING & INFRASTRUCTURE

IO Tech Solutions Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
AI Algorithm Engineer
AI Algorithm Engineer

Cloudnav Ai • Hong Kong

On-site
HKD 700,000 - 1,000,000
AI Agent Security Research Engineer
AI Agent Security Research Engineer

OKX • Hong Kong

On-site
HKD 80,000 - 120,000
Competitive total compensation package
L&D programs and education subsidy
Team building programs
+2
Senior Large Language Model Algorithm Engineer/Expert
Senior Large Language Model Algorithm Engineer/Expert

Binance • Hong Kong

On-site
HKD 861,000 - 1,097,000