Senior Software Engineer

optimum solutions (singapore) pte ltd

Singapore

On-site

SGD 66,960 - 133,920

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optimum Solutions is seeking a skilled AI evaluation engineer for a 6-month contract in Singapore. The role focuses on owning the ACS evaluation workstream, establishing regression tests, and ensuring end-to-end traceability across AI agents and tools.

You will work with EvalsHub-like tech, LangSmith, LangChain, and Temporal to build reusable evaluation pipelines, integrate with CI/CD, and deliver measurable agent-quality outcomes. On-site schedule applies.

Qualifications

  • Minimum 2 years in AI evaluation, tooling, or platform work.
  • Experience with LLM-based evaluation and multi-turn workflows.
  • Knowledge of regression/evaluation pipelines and tooling.
  • Familiarity with embeddings, retrieval, and model benchmarking.
  • Ability to build reusable platform capabilities for AI products.

Responsibilities

  • Own ACS embedded-evaluation workstream and CI/CD integration.
  • Establish reliable regression evaluation for additional teams.
  • Ensure traces are usable for debugging and evaluation.
  • Ship reusable improvements to EvalsHub or SDK/tooling.
  • Reduce manual evaluation with automated datasets and CI integration.
  • Document procedures and onboarding guidance.
  • Collaborate with AI product teams; own quality outcomes.
  • Build golden datasets, evaluation pipelines, and dashboards.
  • Translate product rules into deterministic evaluators and SOPs.
  • Diagnose failures and drive data/tooling improvements.
  • Instrument distributed agent systems with OpenTelemetry.
  • Ensure traces capture inputs, outputs, tools, and results.
  • Develop reusable backend/CLI/frontend capabilities.
  • Integrate evaluations into model release workflows.
  • Run experiments on model, retrieval, and architecture; balance metrics.
  • Lead migration from legacy systems; write designs and demos.
  • Participate in LLMOps operational rotation.

Skills

EvalsHub/LangSmith/LangGraph/LangChain
FastAPI
Temporal
Grafana
Redis
React/TypeScript
LLM-as-judge
Multi-turn evaluation
Agent-trajectory analysis
Text-to-SQL/DSL
RAG & embedding evaluation
Model benchmarking
Platform capabilities building

Tools

EvalsHub
LangSmith
LangGraph/LangChain
FastAPI
Temporal
Grafana
Redis

Job description

This is a 6-month contract role under Optimum Solutions.

This role supports our client, a leading Southeast Asian superapp that offers a wide range of services, including mobility, food delivery, and digital payments.

Work Schedule and Hours: Monday - Friday, 10am - 7pm (inclusive of a 1-hour lunch break)

Key Responsibilities
  • Take ownership of the ACS embedded-evaluation workstream.
  • Establish reliable regression evaluation for at least one additional agent team.
  • Make its distributed traces consistently usable for debugging and automated evaluation.
  • Ship at least one reusable improvement to EvalsHub or its SDK/tooling.
  • Reduce manual evaluation effort through automated datasets, evaluators, and CI integration.
  • Document ownership, operational procedures, and onboarding guidance.
  • Embed with AI product teams and own measurable agent-quality outcomes.
  • Build and maintain golden datasets, regression suites, evaluation pipelines, and quality dashboards.
  • Convert product rules and human review procedures into deterministic, numeric, LLM-as-judge, trajectory, tool-selection, SOP-adherence, and multi-turn evaluators.
  • Diagnose evaluation and production-trace failures, then transform findings into dataset improvements, evaluator changes, or fixes to prompts, tools, skills, and agent workflows.
  • Instrument distributed agent systems using OpenTelemetry, including Temporal workflows and Go/Python services.
  • Ensure traces contain consistent inputs, outputs, tool calls, metadata, feedback, and root-agent results.
  • Develop reusable capabilities across EvalsHub backend, SDK, CLI/plugins, and supporting frontend interfaces.
  • Integrate evaluations into CI/CD and model-release workflows.
  • Run model, retrieval, embedding, and agent-architecture experiments; balance accuracy, latency, and cost.
  • Drive migrations from legacy evaluation and observability systems.
  • Write technical designs and documentation, run demonstrations, and enable engineering teams to use evaluation tooling independently.
  • Participate in the LLMOps operational rotation and support production-quality integrations.
Qualifications
  • Minimum 2 years of relevant experience in one or more of the following areas:
    • EvalsHub, LangSmith, LangGraph/LangChain, FastAPI, Temporal, Grafana, Redis, or similar technologies.
    • React/TypeScript and internal developer‑tooling interfaces.
    • LLM-as-judge, multi-turn evaluation, tool/MCP evaluation, and agent-trajectory analysis.
    • Text-to-SQL/Text-to-DSL, RAG, semantic retrieval, embedding evaluation, or model benchmarking.
    • Building platform capabilities that can be reused across several AI products.
  • Strong analytical and problem-solving skills related to AI systems, evaluation frameworks, and data-driven quality improvements.
  • Excellent project management and cross-department communication abilities.
  • A detail-oriented and metrics-driven approach to continuous improvement.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer(AI & LLM) - 12 Months Contract - up to SGD12k -
Senior Software Engineer(AI & LLM) - 12 Months Contract - up to SGD12k -

MORGAN MCKINLEY PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Senior AI Evaluation Engineer: CI/CD & Quality Metrics
Senior AI Evaluation Engineer: CI/CD & Quality Metrics

optimum solutions (singapore) pte ltd • Singapore

On-site
#EG AI Engineer
#EG AI Engineer

NCS Group • Singapore

On-site
SGD 80,000 - 120,000
Senior AI Evaluation Engineer: LLM Ops & Automation
Senior AI Evaluation Engineer: LLM Ops & Automation

MORGAN MCKINLEY PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
AI Specialist - TEKsystems (Allegis Group Singapore Pte Ltd)
AI Specialist - TEKsystems (Allegis Group Singapore Pte Ltd)

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
AI Tester
AI Tester

NCS PTE. LTD. • Singapore

On-site
SGD 72,000 - 120,000
Senior AI Engineer (Xora Portfolio Company)
Senior AI Engineer (Xora Portfolio Company)

Xora Innovation • Singapore

Hybrid
SGD 192,000 - 269,000
AI Engineer (Managed Services)
AI Engineer (Managed Services)

Avepoint • Singapore

On-site
SGD 90,000 - 120,000
Flexible working hours
Access to high-performance GPU resources
Continued learning and development opportunities
AI Engineer
AI Engineer

NCS PTE. LTD. • Singapore

On-site
SGD 120,000 - 210,000
Full Stack Engineer (Agentic AI)
Full Stack Engineer (Agentic AI)

Synechron • Singapore

On-site
SGD 90,000 - 150,000