AI Engineer (LLM/ Chatbot)

Pantheon Lab Limited

Hong Kong

On-site

HKD 900,000 - 1,300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Pantheon Lab Limited in Hong Kong is seeking an experienced AI Engineer to design, deploy, and optimize production-grade language model systems. You will build applications using both commercial LLM APIs and self-hosted open-source models, implement RAG pipelines, and create end-to-end LLM workflows.

The role blends practical API integration with expert deployment of local models to deliver scalable AI solutions.

Qualifications

  • Solid 2+ years focused on LLM applications or chatbot development.
  • Proven track record of building production LLM applications.
  • Experience integrating and optimizing commercial LLM APIs.
  • Hands-on experience deploying local models in production environments.

Responsibilities

  • Design and implement high-throughput, low-latency serving architectures for LLM applications
  • Build and maintain RAG pipelines and end-to-end LLM workflows
  • Integrate and optimize commercial LLM APIs into production systems
  • Develop prompt engineering techniques and prompt management systems
  • Deploy and serve local open-source language models for specific use cases
  • Optimize local model inference performance through efficient serving frameworks
  • Fine-tune models to improve performance on domain-specific tasks
  • Monitor and troubleshoot production LLM systems to ensure reliability
  • Research and experiment with emerging models and techniques to improve system capabilities
  • Document architectures, best practices, and technical decisions
  • Collaborate with engineering teams to integrate LLM capabilities into products
  • Communicate technical terms and recommendations to stakeholders
  • Build evaluation frameworks to measure model quality, latency, cost, and user satisfaction
  • Design intelligent routing and fallback strategies across multiple LLM providers
  • Scale LLM services to handle production workloads efficiently
  • Implement caching, batching, and request optimization strategies for both APIs and local models

Skills

Python programming
Async/await
Transformer basics
LLM APIs integration
Local model deployment
Docker
FastAPI
Vector databases
MongoDB
Git/GitHub
Hugging Face
Cloud platforms
Kubernetes
LLM fine-tuning
LangChain/LlamaIndex

Tools

LlamaIndex
LangChain
vLLM
MLflow
Phoenix
Langfuse
Weaviate
Chroma
Milvus
Pinecone
TGI
Hugging Face

Job description

About The Role

We are seeking an experienced AI Engineer to manage the design, deployment, and optimization of production-grade language model systems. This role involves building applications using both commercial LLM APIs and self-hosted open-source models, implementing RAG pipelines, and creating end-to-end LLM workflows. The ideal candidate combines practical experience integrating LLM APIs with technical expertise in deploying and optimizing local models.

About The Role

We are seeking an experienced AI Engineer to manage the design, deployment, and optimization of production-grade language model systems. This role involves building applications using both commercial LLM APIs and self-hosted open-source models, implementing RAG pipelines, and creating end-to-end LLM workflows. The ideal candidate combines practical experience integrating LLM APIs with technical expertise in deploying and optimizing local models.

Key Responsibilities
  • Design and implement high-throughput, low-latency serving architectures for LLM applications
  • Build and maintain RAG pipelines and end-to-end LLM workflows
  • Integrate and optimize commercial LLM APIs (OpenAI, Anthropic, Google, etc.) into production systems
  • Develop prompt engineering techniques and prompt management systems
  • Deploy and serve local open-source language models for specific use cases
  • Optimize local model inference performance through efficient serving frameworks
  • Fine-tune models to improve performance on domain-specific tasks (
  • Monitor and troubleshoot production LLM systems to ensure reliability
  • Research and experiment with emerging models and techniques to improve system capabilities
  • Document architectures, best practices, and technical decisions
  • Collaborate with engineering teams to integrate LLM capabilities into products
  • Communicate technical terms and recommendations to stakeholders
  • Build evaluation frameworks to measure model quality, latency, cost, and user satisfaction
  • Design intelligent routing and fallback strategies across multiple LLM providers
  • Scale LLM services to handle production workloads efficiently
  • Implement caching, batching, and request optimization strategies for both APIs and local models
Experience
Required Qualifications
  • Solid 2+ years focused on LLM applications or chatbot development (Candidates with more experience will be considered for a senior role.)
  • Proven track record of building production LLM applications
  • Experience integrating and optimizing commercial LLM APIs
  • Hands-on experience deploying local models in production environments
Technical Skills
  • Strong Python programming with emphasis on async/await patterns and production-quality code
  • Deep understanding of transformer architectures and LLM fundamentals
  • Experience with LLM APIs (OpenAI, Anthropic Claude, Google Gemini, or similar)
  • Familiarity with local open-source models (Qwen, Llama, Mistral, or similar)
  • Experience with RAG implementation using LlamaIndex, LangChain or similar frameworks
  • Proficiency with FastAPI for building high-performance APIs
  • Experience with vector databases (Pinecone, Weaviate, Chroma, Milvus, or similar)
  • Working knowledge of MongoDB or other NoSQL databases
  • Experience with Docker containerization and deployment
  • Good to have hands-on fine-tuning experience (LoRA, QLoRA, full fine-tuning)
  • Familiarity with local model serving frameworks (vLLM, TGI, or similar)
  • Familiarity with LLM workflow tracing and observability frameworks (MLflow, Phoenix, Langfuse, or similar)
  • Familiarity with Hugging Face ecosystem and transformer libraries
  • Experience with cloud platforms (AWS, GCP, or Azure)
  • Proficiency with Git/GitHub and version control workflows
Domain Knowledge
  • Understanding of prompt engineering and optimization techniques
  • Knowledge of LLM evaluation metrics and benchmarking methodologies
  • Experience with cost optimization for LLM applications
  • Familiarity with distributed computing and scaling strategies
  • Understanding of LLM inference optimization (quantization, batching, caching)
Preferred Qualifications
  • Understanding of digital human technologies and multimodal applications
  • Knowledge of MLOps practices and CI/CD for ML systems
  • Experience with Kubernetes for container orchestration
  • Experience with streaming inference and real-time applications
  • Background in function calling and tool use with LLMs
  • Familiarity with RLHF (Reinforcement Learning from Human Feedback)
  • Experience with model distillation and knowledge compression
  • Understanding of distributed training and GPU optimization
  • Experience with multi-agent systems and LLM orchestration
Soft Skills & Communication
  • Excellent English communication skills (written and verbal)
  • Excellent Chinese reading skill
  • Ability to explain complex technical concepts to both technical and non-technical audiences
  • Strong problem-solving and analytical thinking capabilities
  • Self-motivated with ability to work independently and drive projects to completion
  • Collaborative team player who thrives in fast-paced environments
  • Passion for staying current with rapidly evolving LLM technologies
  • Ability to balance research experimentation with production reliability requirements

We can’t wait to hear from you! NOTE: All personal data is collected for recruitment purposes only.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer Intern (RAG/LLM)
AI Engineer Intern (RAG/LLM)

ICC (Hong Kong) Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
LLM Engineer (Data and Optimization)
LLM Engineer (Data and Optimization)

TCL Corporate Research(HK) Co., Ltd • Hong Kong

Hybrid
HKD 900,000 - 1,300,000
AI Engineer - LLM & Production Chatbot Systems
AI Engineer - LLM & Production Chatbot Systems

Pantheon Lab Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
Senior Large Language Model Algorithm Engineer/Expert
Senior Large Language Model Algorithm Engineer/Expert

Binance • Hong Kong

On-site
HKD 861,000 - 1,097,000
AI Engineer (LLM, ML)
AI Engineer (LLM, ML)

PERSOL • Hong Kong

On-site
HKD 600,000 - 900,000
AI Engineer
AI Engineer

ConnectedGroup Limited • Hong Kong

On-site
HKD 600,000 - 900,000
AI Systems Engineer
AI Systems Engineer

SENSLAB HK LIMITED • Hong Kong

On-site
HKD 360,000 - 600,000
AI Engineer
AI Engineer

AXG (Solowin Holdings) • Hong Kong Island

On-site
HKD 1,000,000 - 1,800,000
Quant Research Engineer
Quant Research Engineer

Millennium • Hong Kong

On-site
HKD 900,000 - 1,300,000
AI Engineer (Product)
AI Engineer (Product)

FinCatch Limited • Hong Kong

On-site
HKD 500,000 - 1,000,000