AI Engineer — LLM Systems & Model Training

Zebra Strategic Outsource Solution Ltd

Hong Kong

On-site

HKD 600,000 - 900,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Zebra Strategic Outsource Solution Ltd is hiring for a hands-on ML/LLM engineer to architect modular system prompts, build long-term memory, and own response management across English, Cantonese and Mandarin. You will select models, lead data pipelines, and train/open-weight LLMs while optimizing GPU deployments for cost and speed.

Candidates should have strong Python/PyTorch skills and exposure to Hugging Face tools.

Qualifications

  • 4+ years in ML or software engineering with 1+ years shipping LLM products to real users.
  • Deep, hands-on prompt engineering — system-prompt design, few-shot, structured output, tool calling, prompt testing.
  • You have built memory or RAG systems in production.
  • You have fine-tuned open-weight LLMs and built the dataset yourself.
  • Strong Python and PyTorch; comfortable with Hugging Face, distributed training and cloud GPUs.
  • Experimentation rigor — you can describe an eval or test that failed and what you changed.
  • Startup mindset — you ship fast, own outcomes and work directly with the founders.
  • English and Chinese required; Japanese is a strong plus.

Responsibilities

  • Architect a modular system-prompt framework — character identity, voice and style guides, world rules and safety layers per character/context.
  • Build long-term memory — summarization, retrieval (RAG / vector DB), user-profile extraction, recall at the right moment.
  • Own response management — persona consistency, tone and length control, safety refusals in character.
  • Select and orchestrate models — frontier APIs and open-weight bases based on quality, license, cost.
  • Lead the training-data pipeline — sourcing, consent, cleaning, de-duplication, PII removal, labelling and versioning.
  • Generate and curate synthetic dialogue data; write annotation guidelines and manage labelling vendors.
  • Train our own models — continued pre-training, SFT, LoRA and preference tuning (DPO / RLHF).
  • Distil, quantize and deploy on GPU inference stacks (e.g. vLLM, TensorRT-LLM); manage compute budget.
  • Build evaluation across tracks — golden sets, LLM-as-judge, persona-drift and safety red-teaming.
  • Ship in English, Cantonese and Mandarin.

Skills

ML engineering
Prompt engineering
Memory / RAG systems
Open-weight LLM fine-tuning
Python
PyTorch
Hugging Face
Distributed training
Cloud GPUs
Experimentation
Startup mindset
English and Chinese

Tools

Hugging Face
TensorRT-LLM
vLLM
distributed training

Job description

You own the brain behind every character. You run two tracks in parallel: the LLM system that makes conversations feel consistent, personal and alive today, and our own models trained on data you source and shape. Each feeds the other — production conversations become training data, and trained models replace prompts where they win on quality or cost. This is the first dedicated hire on the model layer; the architecture is yours to define.

Key responsibilities

  • Architect a modular system-prompt framework — character identity, voice and style guides, world rules and safety layers composed per character and per context, versioned and testable.
  • Build long-term memory — what to remember, summarization, retrieval (RAG / vector DB), user-profile extraction, recall at the right moment, forgetting and correction.
  • Own response management — persona consistency, tone and length control, emotional-context handling, anti-repetition, multi-turn coherence, guardrails and refusals that stay in character.
  • Select and orchestrate models — frontier APIs and open-weight bases chosen on quality, license, size, language coverage and serving cost; routing, fallback, streaming, latency and cost per conversation.
  • Lead the training-data pipeline — sourcing, consent, cleaning, de-duplication, PII removal, quality filtering, labelling and dataset versioning, including turning live conversations into training data.
  • Generate and curate synthetic dialogue data; write annotation guidelines and manage labelling vendors.
  • Train our own models — continued pre-training, SFT, LoRA and preference tuning (DPO / RLHF) for character- and domain-specialized models.
  • Distil, quantize and deploy on GPU inference stacks (e.g. vLLM, TensorRT-LLM); own the compute budget.
  • Build evaluation across both tracks — golden sets, LLM-as-judge, persona-drift and safety red-teaming, A/B tests tied to retention and engagement — so prompt and model versions compete on the same metrics.
  • Ship in English, Cantonese (including colloquial written Cantonese) and Mandarin.

About you

  • 4+ years in ML or software engineering, with 1+ years shipping LLM products to real users.
  • Deep, hands-on prompt engineering — system-prompt design, few-shot, structured output, tool calling, prompt testing.
  • You have built memory or RAG systems in production.
  • You have fine-tuned open-weight LLMs (DeepSeek, Qwen, Llama, Mistral, Gemma or similar) and built the dataset yourself.
  • Strong Python and PyTorch; comfortable with Hugging Face, distributed training and cloud GPUs.
  • Experimentation rigor — you can tell us about an eval or test that failed and what you changed.
  • Startup mindset — you ship fast, own outcomes and work directly with the founders.
  • English and Chinese required; Japanese is a strong plus.

About us

We are a funding-backed team working at the frontier of generative AI, building a next-generation AI social gaming platform — where AI-driven characters, stories, and interactive experiences bring people together in entirely new ways. Our characters hold long, personal, multi-session conversations with every player. You will work directly with the founding team and see your work live with real users.

Unlock job insights

Hirer responsiveness Salary match Number of applicants

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Engineer (LLM/ Chatbot)
AI Engineer (LLM/ Chatbot)

Pantheon Lab Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
AI Prompter (Prompt Engineer)
AI Prompter (Prompt Engineer)

Occasions Asia Pacific 天機亞太集團 • Hong Kong

On-site
HKD 480,000 - 800,000
5 days work-week
Medical plan
Dental insurance
AI Engineer
AI Engineer

Hyphen Connect Limited • Hong Kong

On-site
HKD 900,000 - 1,200,000
Forward Deployed Engineer, Hong Kong
Forward Deployed Engineer, Hong Kong

Telnyx • Hong Kong

On-site
HKD 900,000 - 1,300,000
NLP & Machine Learning Engineer
NLP & Machine Learning Engineer

Datago Technology Limited • Hong Kong

On-site
HKD 420,000 - 640,000
Annual leave
Group medical insurance
Visa assistance
Staff AI Engineer
Staff AI Engineer

Wati Dot I O • Hong Kong

On-site
HKD 900,000 - 1,500,000
Direct access to founding team
Ownership of AI roadmap
AI Engineer
AI Engineer

NCS Pte Ltd • Hong Kong

On-site
HKD 900,000 - 1,200,000
Character AI LLM Architect & Model Training Lead
Character AI LLM Architect & Model Training Lead

Zebra Strategic Outsource Solution Ltd • Hong Kong

On-site
HKD 600,000 - 900,000
UI/UX Lead
UI/UX Lead

Hyphen Connect • Hong Kong

On-site
HKD 700,000 - 900,000
Vice President, AI Strategy & Transformation
Vice President, AI Strategy & Transformation

OKX • Hong Kong

On-site
HKD 3,000,000 - 5,400,000
Competitive total compensation package
L&D programs and education subsidy
Comprehensive healthcare schemes
+3