We are looking for a Senior AI Systems Engineer with deep experience in multi-model systems, model routing, AI evaluation, and production AI infrastructure in Singapore or Hong Kong. The role will focus on designing systems that can compare, select, and operate multiple foundation models reliably in production.
The ideal candidate should be able to define evaluation standards, translate evaluation results into routing strategies, and build the supporting platform required to optimize model quality, latency, cost, and reliability.
- Design and implement model-routing strategies across multiple foundation models based on task type, complexity, quality requirements, latency, cost, and operational risk.
- Develop dynamic model-selection, arbitration, escalation, fallback, and ensemble mechanisms.
- Define routing policies and decision criteria for different task categories and production scenarios.
- Continuously improve routing performance using evaluation results, production data, and controlled experimentation.
- Distinguish and optimize for model capability rather than relying solely on availability-based failover or static rule-based routing.
- Define task taxonomies, evaluation dimensions, scoring criteria, and acceptance thresholds for different AI use cases.
- Design and maintain benchmark suites, golden datasets, annotation standards, and regression test sets.
- Build automated evaluation pipelines using deterministic checks, LLM-as-a-Judge, rubric-based scoring, pairwise comparison, and human evaluation where appropriate.
- Validate evaluation methods against human judgments and monitor judge consistency, bias, and drift.
- Establish continuous evaluation and regression mechanisms for model, prompt, data, and routing-policy changes.
- Ensure evaluation signals are reliable enough to support model comparison and routing decisions.
- Design and build scalable infrastructure for integrating and serving multiple proprietary and open-source models.
- Develop model gateways, unified APIs, version-management mechanisms, and routing infrastructure.
- Support model lifecycle management, traffic governance, rollout, rollback, and model replacement.
- Improve the reliability, scalability, and operational efficiency of multi-model production systems.
- Define measurable optimization objectives across model quality, inference cost, latency, reliability, and resource utilization.
- Develop systematic methods to balance competing objectives rather than optimizing a single metric in isolation.
- Compare routing strategies against fixed-model baselines and quantify their operational and performance benefits.
- Use offline evaluation, online experimentation, and production feedback to improve routing policies continuously.
- Build production-grade AI services with strong standards for reliability, security, testing, and maintainability.
- Implement monitoring, logging, tracing, failure analysis, and performance diagnostics for model calls and routing decisions.
- Establish data feedback loops to capture model performance, routing outcomes, failure cases, and human-review results.
- Support incident investigation, production debugging, regression analysis, and system recovery.
- Hands-on experience building production AI, machine-learning, or distributed systems.
- Experience designing evaluation frameworks, benchmark datasets, scoring methodologies, regression pipelines, or continuous evaluation systems.
- Strong understanding of the trade-offs among model quality, latency, inference cost, reliability, and operational risk.
- Experience integrating and operating multiple foundation models, including third-party APIs, self-hosted models, or open-source models.
- Ability to translate ambiguous requirements into measurable evaluation criteria, technical designs, and production systems.
- Strong analytical and problem-solving skills, with the ability to validate technical decisions using data and experiments.
- Experience with LLM-as-a-Judge, human evaluation, judge calibration, preference evaluation, or benchmark development.
- Experience with model serving, inference optimization, model gateways, observability platforms, or LLMOps systems.
- Experience with reinforcement learning, contextual bandits, ranking, recommendation systems, search, or information retrieval.
- Experience with AI safety, red teaming, model governance, or evaluation of high-risk AI systems.
- You think in terms of systems, objectives, and measurable outcomes rather than individual prompts or isolated model outputs.
- You can distinguish true model routing from static rules, failover, or load balancing.
- You understand that reliable routing depends on reliable evaluation.
- You are comfortable owning both technical strategy and production implementation.