Edge Inference & Systems Optimization Engineer

奇瑞全球創新(香港)有限公司

Hong Kong Island

On-site

HKD 900,000 - 1,300,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

奇瑞全球創新(香港)有限公司 is seeking a senior ML inference engineer to design and build high-performance inference serving infrastructure for model deployment across edge devices and heterogeneous hardware.

You will optimize latency, throughput, and cost through end-to-end engineering, including dynamic batching, memory management, and GPU kernel tuning, while collaborating with algorithm and research teams to ensure deployment readiness.

Qualifications

  • End-to-end engineering from model to inference system and edge deployment.
  • Benchmark and profile performance across diverse hardware including GPUs and edge devices.
  • Collaborate with algorithm and research teams to ensure deployment readiness from day one.
  • Explore next-generation inference paradigms such as speculative decoding and mixture-of-experts routing.

Responsibilities

  • Design and build high-performance inference serving infrastructure, including request scheduling, dynamic batching, and memory management.
  • Benchmark and profile model performance across diverse hardware; identify bottlenecks and drive co-optimization.
  • Collaborate with algorithm and research teams to ensure new models are deployment-ready from day one.
  • Explore and implement next-generation inference paradigms: speculative decoding, mixture-of-experts routing, and sparse attention.

Skills

End-to-end engineering
Model inference optimization
Transformer architectures
CUDA programming
GPU kernel optimization
System performance analysis
Edge deployment
Collaboration with research teams
Memory architecture understanding
Edge hardware collaboration

Education

Bachelor’s degree or above in Computer Science, Electronic Engineering, or related fields

Tools

CUDA
TensorRT-LLM
vLLM
DeepSpeed-Inference
Ascend CANN

Job description

奇瑞全球創新(香港)有限公司 is seeking a senior ML inference engineer to design and build high-performance inference serving infrastructure for model deployment across edge devices and heterogeneous hardware.

You will optimize latency, throughput, and cost through end-to-end engineering, including dynamic batching, memory management, and GPU kernel tuning, while collaborating with algorithm and research teams to ensure deployment readiness.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference & System Optimization Engineer (Experienced )
Inference & System Optimization Engineer (Experienced )

奇瑞全球創新(香港)有限公司 • Hong Kong Island

On-site
HKD 900,000 - 1,300,000
AI Engineer: Edge ML & Scalable Models
AI Engineer: Edge ML & Scalable Models

Novos Technology Limited • Hong Kong

On-site
HKD 900,000 - 1,500,000
ML Engineer: Real-Time Inference & Large-Scale Training
ML Engineer: Real-Time Inference & Large-Scale Training

ittihad medical centre • Hong Kong

On-site
HKD 500,000 - 700,000
ML Engineer: Real-Time Inference & Large-Scale Training
ML Engineer: Real-Time Inference & Large-Scale Training

IMC B.V. • Hong Kong

On-site
HKD 400,000 - 600,000
Production AI Engineer: Build & Deploy ML Inference
Production AI Engineer: Build & Deploy ML Inference

香港亞方科技有限公司 • Hong Kong

On-site
HKD 600,000 - 1,000,000
Distributed ML Engineer: Real-Time Inference & GPU Training
Distributed ML Engineer: Real-Time Inference & GPU Training

IMC Trading • Hong Kong

On-site
HKD 700,000 - 1,100,000
LLM Algorithm Architect — Lead AI Innovation
LLM Algorithm Architect — Lead AI Innovation

奇瑞全球創新(香港)有限公司 • Hong Kong Island

On-site
HKD 900,000 - 1,500,000
Senior ML Research Engineer - Real-Time AI Platform
Senior ML Research Engineer - Real-Time AI Platform

Sanderson-Ikas • Hong Kong

On-site
HKD 900,000 - 1,500,000
Applied ML Engineer: Turn Research into Production AI
Applied ML Engineer: Turn Research into Production AI

Sentient Labs • Hong Kong

On-site
HKD 700,000 - 1,000,000
CUDA Kernel Architect for Low-Latency Inference
CUDA Kernel Architect for Low-Latency Inference

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 900,000 - 1,300,000