Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs

Irvine (CA)

Hybrid

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitus Labs is seeking an AI Engineer to drive agent systems, memory management, and RAG pipelines for our in-game personalization platform. You will contribute to model deployment, inference optimization, and end-to-end production readiness.

This role focuses on building and operating scalable, low-latency AI services with Docker/Kubernetes, LangGraph, and vector databases, while collaborating with researchers and backend engineers to ship production-ready systems.

Qualifications

  • Bachelor's or Master's in CS/ML or related field.
  • 3+ years of industry experience in ML engineering or related roles.
  • Proven ability to build low-latency production services in Go, Rust, C++, or equivalents.
  • Strong PyTorch and Hugging Face experience for production LLMs and agents.
  • Experience with real-time inference systems and vector databases.

Responsibilities

  • Design, build, and optimize LLM-powered agents with planning, tool use, and multi-step reasoning.
  • Architect memory systems and manage context/session state.
  • Build and optimize RAG pipelines for relevance and freshness.
  • Operate vector-store infra (pgvector, Milvus, Qdrant, Weaviate).
  • Define evaluation methods for agents and prompts; improve latency and reliability.
  • Develop production inference services with low latency and high concurrency.

Skills

Go / Rust / C++
PyTorch / HuggingFace
LLM / agent apps
Docker / Kubernetes
Cloud platforms (AWS / GCP / Azure)
Real-time inference
Production ML systems

Education

Bachelor's or Master's in CS/ML

Tools

LangGraph
LlamaIndex
pgvector
Milvus
Qdrant
Weaviate

Job description

We're an online gaming company using AI to power and personalize player experiences. This role sits within the AI Engineering team, which is responsible for taking AI capabilities into production. This role focuses primarily on agent systems, with model deployment and inference engineering as a secondary responsibility.

We do not train foundation models from scratch. Our focus is on production AI systems, model adaptation, inference optimization, and agentic applications.

Responsibilities
Agent Systems — Primary
  • Design, build, and optimize LLM-powered agents, including planning, tool use, workflow orchestration, and multi-step reasoning
  • Architect memory systems, including short-term memory, long-term memory, context management, and session state
  • Build and optimize RAG pipelines for relevance, grounding, freshness, and retrieval quality
  • Design and operate vector-store infrastructure (e.g., pgvector, Milvus, Qdrant, Weaviate)
  • Define evaluation methodologies for agents, prompts, and workflows
  • Optimize end-to-end agent quality, latency, reliability, and operating cost
Model Deployment & Inference — Secondary
  • Build and operate production inference services that are low-latency, high-concurrency, and highly reliable
  • Serve online-learning models (e.g., contextual bandits and reinforcement learning policies) with real-time inference and online parameter or weight updates
  • Deploy and optimize AI inference systems for latency, throughput, reliability, and resource efficiency
  • Analyze and resolve inference-serving bottlenecks
  • Support deployment and serving of recommendation, ranking, and reinforcement learning models developed by research scientists
  • Apply lightweight model adaptation techniques (e.g., LoRA, QLoRA, PEFT) when appropriate for domain-specific requirements
MLOps — Supporting Both
  • Build and maintain deployment pipelines, observability systems, and tracing infrastructure for agents and serving endpoints
  • Monitor quality regression, performance degradation, and model drift
  • Maintain version control for models, prompts, datasets, and agent configurations
  • Contribute to automated validation, testing, and CI/CD workflows for AI systems
  • Partner with research scientists, backend engineers, and data scientists to integrate AI systems into production products
  • Document systems, best practices, and internal tooling
  • Contribute to engineering standards and operational excellence across AI initiatives
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Machine Learning, or a related field
  • 3+ years of industry experience in Machine Learning Engineering or related roles
  • Strong software and systems engineering experience, including building low-latency, reliable production services in languages such as Go, Rust, C++, or equivalent
  • Experience building or supporting real-time inference systems for recommendation, ranking, contextual bandits, reinforcement learning, or similar adaptive machine learning applications Strong experience with PyTorch and the Hugging Face ecosystem
  • Experience building production LLM or agent applications (e.g., LangGraph, LlamaIndex, or equivalent frameworks)
  • Hands-on experience with RAG systems, embeddings, and vector databases
  • Experience evaluating and monitoring LLM or agent systems in production
  • Experience deploying and optimizing production machine learning or LLM systems
  • Understanding of inference runtime behavior, resource utilization, latency optimization, and production serving performance
  • Experience with Docker and Kubernetes
  • Experience with cloud platforms such as AWS, GCP, or Azure
Preferred / Nice to Have
  • Experience fine-tuning open-weight LLMs using LoRA, QLoRA, PEFT, or related approaches
  • Familiarity with the underlying algorithms used in recommender systems, ranking systems, contextual bandits, or reinforcement learning
  • Experience with custom GPU kernel development using CUDA or OpenAI Triton
  • Experience with graph-level optimization and low-level inference performance tuning
  • Experience with large-scale distributed training (e.g., FSDP, DeepSpeed, multi-GPU workloads)
  • Experience deploying models to edge environments using TFLite, CoreML, or NPU accelerators
  • Strong understanding of CI/CD principles and deployment workflows
  • Background in gaming, gaming AI, or player personalization systems
  • Experience with distributed systems, Spark, Hadoop, or large-scale data infrastructure
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Engineer
Senior AI Engineer

7AI • Boston (MA)

On-site
USD 140,000 - 200,000
AI/ML Engineer
AI/ML Engineer

RiskForce • Northern (KY)

Hybrid
USD 120,000 - 155,000
Machine Learning Engineer
Machine Learning Engineer

Orbien LLC • Maryland

On-site
USD 150,000 - 230,000
AI Engineer, Multimodal LLMs
AI Engineer, Multimodal LLMs

eloquentai • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Algo Eng - LLM/VLM Mandarin Required
AI Algo Eng - LLM/VLM Mandarin Required

Fuku • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Algo Eng – LLM/VLM (Mandarin Required)
AI Algo Eng – LLM/VLM (Mandarin Required)

Applied Intelligence Consulting (Singapore) • San Francisco (CA)

On-site
USD 180,000 - 280,000
Lead Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Lead Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI

Nutanix • Santa Clara (CA)

On-site
USD 120,000 - 150,000
Family medical coverage
Vision coverage
Dental coverage
+1
Principal AI Engineer
Principal AI Engineer

Stellantis NV • Auburn (AL)

On-site
USD 150,000 - 230,000
Principal AI Engineer
Principal AI Engineer

Stellantis • Auburn Hills (MI)

On-site
USD 180,000 - 280,000
AI Engineer
AI Engineer

Kaleidoscope Innovation • Fort Worth (TX)

On-site
USD 140,000 - 190,000