Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
奇瑞全球創新(香港)有限公司 is seeking a senior ML inference engineer to design and build high-performance inference serving infrastructure for model deployment across edge devices and heterogeneous hardware.
You will optimize latency, throughput, and cost through end-to-end engineering, including dynamic batching, memory management, and GPU kernel tuning, while collaborating with algorithm and research teams to ensure deployment readiness.
End-to-end engineering from model to inference system, hardware platform, and edge-side deployment
Continuous optimization of inference cost, latency, throughput, and stability
Key technologies: quantization, pruning, distillation, KV Cache, Continuous Batching, Paged Attention, Speculative Decoding, CUDA/Ascend/PPU operator optimization, multi-GPU distributed inference, model compilation, and cloud-edge collaboration
Design and build high-performance inference serving infrastructure, including request scheduling, dynamic batching, memory management, and fault tolerance
Benchmark and profile model performance across diverse hardware (NVIDIA GPU, Huawei Ascend, edge devices), identify bottlenecks, and drive hardware-software co-optimization
Collaborate with algorithm and research teams to ensure new models are deployment-ready from day one; provide early feedback on architecture decisions that affect inference efficiency
Explore and implement next-generation inference paradigms: speculative decoding, mixture-of-experts routing, sparse attention, and model-hardware co-design
Hands-on experience in large model inference efficiency optimization, with a proven track record of measurable improvements in latency, throughput, or cost
Deep understanding of transformer architecture, attention mechanisms, and memory-bound vs. compute-bound workload characteristics
Proficiency in CUDA programming and GPU kernel optimization; experience with Ascend CANN or other domestic AI chip stacks is a strong plus
Solid experience with mainstream inference frameworks: vLLM, TensorRT-LLM, SGLang, DeepSpeed-Inference, or similar
Familiarity with model compression techniques (quantization INT4/INT8/FP8, pruning, knowledge distillation) and their trade-offs in accuracy vs. efficiency
Strong systems thinking: able to reason about end-to-end performance from model architecture through serving infrastructure to hardware constraints
Experience deploying models on edge devices or heterogeneous computing environments is preferred
Bachelor’s degree or above in Computer Science, Electronic Engineering, or related fields; 3+ years of relevant industry experience