Inference & System Optimization Engineer (Experienced )

奇瑞全球創新(香港)有限公司

Hong Kong Island

On-site

HKD 900,000 - 1,300,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

奇瑞全球創新(香港)有限公司 is seeking a senior ML inference engineer to design and build high-performance inference serving infrastructure for model deployment across edge devices and heterogeneous hardware.

You will optimize latency, throughput, and cost through end-to-end engineering, including dynamic batching, memory management, and GPU kernel tuning, while collaborating with algorithm and research teams to ensure deployment readiness.

Qualifications

  • End-to-end engineering from model to inference system and edge deployment.
  • Benchmark and profile performance across diverse hardware including GPUs and edge devices.
  • Collaborate with algorithm and research teams to ensure deployment readiness from day one.
  • Explore next-generation inference paradigms such as speculative decoding and mixture-of-experts routing.

Responsibilities

  • Design and build high-performance inference serving infrastructure, including request scheduling, dynamic batching, and memory management.
  • Benchmark and profile model performance across diverse hardware; identify bottlenecks and drive co-optimization.
  • Collaborate with algorithm and research teams to ensure new models are deployment-ready from day one.
  • Explore and implement next-generation inference paradigms: speculative decoding, mixture-of-experts routing, and sparse attention.

Skills

End-to-end engineering
Model inference optimization
Transformer architectures
CUDA programming
GPU kernel optimization
System performance analysis
Edge deployment
Collaboration with research teams
Memory architecture understanding
Edge hardware collaboration

Education

Bachelor’s degree or above in Computer Science, Electronic Engineering, or related fields

Tools

CUDA
TensorRT-LLM
vLLM
DeepSpeed-Inference
Ascend CANN

Job description

End-to-end engineering from model to inference system, hardware platform, and edge-side deployment

Continuous optimization of inference cost, latency, throughput, and stability

Key technologies: quantization, pruning, distillation, KV Cache, Continuous Batching, Paged Attention, Speculative Decoding, CUDA/Ascend/PPU operator optimization, multi-GPU distributed inference, model compilation, and cloud-edge collaboration

Design and build high-performance inference serving infrastructure, including request scheduling, dynamic batching, memory management, and fault tolerance

Benchmark and profile model performance across diverse hardware (NVIDIA GPU, Huawei Ascend, edge devices), identify bottlenecks, and drive hardware-software co-optimization

Collaborate with algorithm and research teams to ensure new models are deployment-ready from day one; provide early feedback on architecture decisions that affect inference efficiency

Explore and implement next-generation inference paradigms: speculative decoding, mixture-of-experts routing, sparse attention, and model-hardware co-design

Hands-on experience in large model inference efficiency optimization, with a proven track record of measurable improvements in latency, throughput, or cost

Deep understanding of transformer architecture, attention mechanisms, and memory-bound vs. compute-bound workload characteristics

Proficiency in CUDA programming and GPU kernel optimization; experience with Ascend CANN or other domestic AI chip stacks is a strong plus

Solid experience with mainstream inference frameworks: vLLM, TensorRT-LLM, SGLang, DeepSpeed-Inference, or similar

Familiarity with model compression techniques (quantization INT4/INT8/FP8, pruning, knowledge distillation) and their trade-offs in accuracy vs. efficiency

Strong systems thinking: able to reason about end-to-end performance from model architecture through serving infrastructure to hardware constraints

Experience deploying models on edge devices or heterogeneous computing environments is preferred

Bachelor’s degree or above in Computer Science, Electronic Engineering, or related fields; 3+ years of relevant industry experience

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Edge Inference & Systems Optimization Engineer
Edge Inference & Systems Optimization Engineer

奇瑞全球創新(香港)有限公司 • Hong Kong Island

On-site
HKD 900,000 - 1,300,000
LLM Engineer (Data and Optimization)
LLM Engineer (Data and Optimization)

TCL Corporate Research(HK) Co., Ltd • Hong Kong

On-site
HKD 900,000 - 1,300,000
Machine Learning Engineer
Machine Learning Engineer

IMC B.V. • Hong Kong

On-site
HKD 400,000 - 600,000
CUDA Kernel Architect for Low-Latency Inference
CUDA Kernel Architect for Low-Latency Inference

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 900,000 - 1,300,000
Machine Learning Engineer
Machine Learning Engineer

ittihad medical centre • Hong Kong

On-site
HKD 500,000 - 700,000
GPU Developer
GPU Developer

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 900,000 - 1,300,000
AI Engineer (LLM/ Chatbot)
AI Engineer (LLM/ Chatbot)

Pantheon Lab Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
CUDA Kernel Engineer for Low-Latency Inference
CUDA Kernel Engineer for Low-Latency Inference

SIG • Hong Kong

On-site
HKD 900,000 - 1,300,000
GPU Developer
GPU Developer

SIG • Hong Kong

On-site
HKD 900,000 - 1,300,000
Principal engineer – GPU Memory Systems & Scale-Up/Out Systems
Principal engineer – GPU Memory Systems & Scale-Up/Out Systems

METAVERSE COMPUTING LIMITED • Hong Kong

On-site
HKD 1,200,000 - 1,600,000