You own the brain behind every character. You run two tracks in parallel: the LLM system that makes conversations feel consistent, personal and alive today, and our own models trained on data you source and shape. Each feeds the other — production conversations become training data, and trained models replace prompts where they win on quality or cost. This is the first dedicated hire on the model layer; the architecture is yours to define.
Key responsibilities
- Architect a modular system-prompt framework — character identity, voice and style guides, world rules and safety layers composed per character and per context, versioned and testable.
- Build long-term memory — what to remember, summarization, retrieval (RAG / vector DB), user-profile extraction, recall at the right moment, forgetting and correction.
- Own response management — persona consistency, tone and length control, emotional-context handling, anti-repetition, multi-turn coherence, guardrails and refusals that stay in character.
- Select and orchestrate models — frontier APIs and open-weight bases chosen on quality, license, size, language coverage and serving cost; routing, fallback, streaming, latency and cost per conversation.
- Lead the training-data pipeline — sourcing, consent, cleaning, de-duplication, PII removal, quality filtering, labelling and dataset versioning, including turning live conversations into training data.
- Generate and curate synthetic dialogue data; write annotation guidelines and manage labelling vendors.
- Train our own models — continued pre-training, SFT, LoRA and preference tuning (DPO / RLHF) for character- and domain-specialized models.
- Distil, quantize and deploy on GPU inference stacks (e.g. vLLM, TensorRT-LLM); own the compute budget.
- Build evaluation across both tracks — golden sets, LLM-as-judge, persona-drift and safety red-teaming, A/B tests tied to retention and engagement — so prompt and model versions compete on the same metrics.
- Ship in English, Cantonese (including colloquial written Cantonese) and Mandarin.
About you
- 4+ years in ML or software engineering, with 1+ years shipping LLM products to real users.
- Deep, hands-on prompt engineering — system-prompt design, few-shot, structured output, tool calling, prompt testing.
- You have built memory or RAG systems in production.
- You have fine-tuned open-weight LLMs (DeepSeek, Qwen, Llama, Mistral, Gemma or similar) and built the dataset yourself.
- Strong Python and PyTorch; comfortable with Hugging Face, distributed training and cloud GPUs.
- Experimentation rigor — you can tell us about an eval or test that failed and what you changed.
- Startup mindset — you ship fast, own outcomes and work directly with the founders.
- English and Chinese required; Japanese is a strong plus.
About us
We are a funding-backed team working at the frontier of generative AI, building a next-generation AI social gaming platform — where AI-driven characters, stories, and interactive experiences bring people together in entirely new ways. Our characters hold long, personal, multi-session conversations with every player. You will work directly with the founding team and see your work live with real users.
Unlock job insights
Hirer responsiveness Salary match Number of applicants