Senior AI Harness Engineer

AVENSYS CONSULTING PTE. LTD.

Singapore

On-site

SGD 150,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Avensys Consulting Pte. Ltd. is seeking a Senior AI Harness Engineer to design, build, and maintain frameworks, evaluation systems, testing infrastructure, and tooling for production-ready AI/LLM applications.

The role focuses on creating AI harnesses that enable systematic testing, benchmarking, observability, and continuous improvement of Generative AI and Agentic AI systems. Strong Python and AI evaluation expertise are required for enterprise-scale projects.

Qualifications

  • 8+ years of software/engineering experience.
  • Experience in AI/ML or Generative AI engineering is required.
  • Strong hands-on experience with Python and AI/LLM technologies.
  • Proven experience developing AI evaluation, testing, automation, or LLMOps solutions.
  • Experience with enterprise-scale AI applications and production deployments.

Responsibilities

  • Design and implement AI/LLM evaluation and testing harnesses for Generative AI and Agentic AI applications.
  • Build reusable frameworks for model evaluation, prompts, regression testing, benchmarking, and performance validation.
  • Develop automated test suites to assess accuracy, relevance, groundedness, toxicity, and latency of AI systems.
  • Create testing frameworks for multi-agent workflows, RAG pipelines, and tool-calling apps.
  • Develop Python-based AI engineering frameworks, utilities, SDKs, and automation tools.
  • Implement MLOps/LLMOps pipelines for model, prompt, dataset, and evaluation lifecycle management.

Skills

Python
REST APIs
JSON
Asynchronous programming
SDK development
Microservices
Git
GitHub
GitLab
Bitbucket
Software engineering best practices

Education

Bachelor's degree in Computer Science

Tools

LangChain
LangGraph
LlamaIndex
Semantic Kernel
AutoGen
CrewAI
OpenAI Agents SDK
MCP (Model Context Protocol)

Job description

Avensys is a reputed global IT professional services company headquartered in Singapore. Our service spectrum includes enterprise solution consulting, business intelligence, business process automation and managed services. Given our decade of success we have evolved to become one of the top trusted providers in Singapore and service a client base across banking and financial services, insurance, information technology, healthcare, retail, and supply chain.

We are currently looking to hire Senior AI Harness Engineer.

This is an exciting opportunity to expand your skill set, achieve job satisfaction and work-life balance. More details as below.

Job Role
Job Description
Senior AI Harness Engineer
Position Overview

We are looking for a Senior AI Harness Engineer to design, build, and maintain the engineering frameworks, evaluation systems, testing infrastructure, and tooling required to develop reliable, scalable, and production-ready AI/LLM applications. The role will focus on building AI harnesses that enable systematic testing, evaluation, benchmarking, observability, and continuous improvement of Generative AI and Agentic AI systems.

The ideal candidate will have strong experience in Python, LLMs, Generative AI, AI agents, evaluation frameworks, RAG, prompt engineering, API integration, cloud platforms, and MLOps/LLMOps, with the ability to build robust engineering solutions around AI models.

Key Responsibilities
  • Design and develop AI/LLM evaluation and testing harnesses for Generative AI and Agentic AI applications.
  • Build reusable frameworks for model evaluation, prompt testing, regression testing, benchmarking, and performance validation.
  • Develop automated test suites to evaluate accuracy, relevance, groundedness, hallucination, toxicity, safety, latency, cost, and response quality.
  • Create testing frameworks for LLM-based agents, multi-agent workflows, RAG pipelines, and tool-calling applications.
  • Develop and maintain Python-based AI engineering frameworks, utilities, SDKs, and automation tools.
  • Implement LLMOps/MLOps pipelines for model, prompt, dataset, and evaluation lifecycle management.
  • Build automated CI/CD pipelines for AI applications, including model and prompt regression testing.
  • Integrate AI evaluation tools and frameworks such as MLflow, LangSmith, DeepEval, Ragas, Azure AI evaluation, OpenAI evaluation frameworks, or equivalent technologies.
  • Design evaluation datasets, test cases, golden datasets, benchmark datasets, and synthetic test data.
  • Develop automated mechanisms for prompt/version management and experiment tracking.
  • Implement observability and monitoring for AI applications, including LLM traces, token usage, latency, errors, model performance, and cost.
  • Develop harnesses for testing RAG systems, including retrieval quality, chunking strategies, embeddings, vector search, reranking, and grounded responses.
  • Build evaluation capabilities for function calling, API/tool invocation, MCP-based integrations, and agentic workflows.
  • Work with AI/ML engineers and application teams to identify failure patterns and improve model/application performance.
  • Establish engineering standards for AI quality, reliability, security, scalability, and responsible AI.
  • Troubleshoot complex issues across application code, LLM APIs, vector databases, cloud services, and AI infrastructure.
  • Build dashboards and reports to communicate AI evaluation and quality metrics to technical stakeholders.
  • Mentor junior engineers and contribute to architecture and technical design decisions.
Required Technical Skills
Programming
  • Strong hands-on experience with Python.
  • Good knowledge of REST APIs, JSON, asynchronous programming, SDK development, and microservices.
  • Experience with Git, GitHub/GitLab/Bitbucket, code reviews, and software engineering best practices.
Generative AI & LLM
  • Strong understanding of LLMs and Generative AI application architecture.
  • Experience with models such as OpenAI, Azure OpenAI, Anthropic, Gemini, Llama, Mistral, or equivalent.
Strong knowledge of:
  • o Prompt Engineering
  • o RAG
  • o Embeddings
  • o Vector databases
  • o Function/Tool Calling
  • o Structured outputs
  • o Agentic AI
  • o LLM APIs
  • o Context management
  • o Model evaluation
AI/Agent Frameworks
Hands-on experience with one or more:
  • LangChain
  • LangGraph
  • LlamaIndex
  • Semantic Kernel
  • AutoGen
  • CrewAI
  • OpenAI Agents SDK
  • MCP (Model Context Protocol)
AI Evaluation & Testing
Experience with AI/LLM evaluation frameworks such as:
  • Ragas
  • DeepEval
  • LangSmith
  • MLflow
  • Azure AI Evaluation
  • OpenAI evaluation frameworks
  • Custom LLM evaluation frameworks
Knowledge of evaluation metrics including:
  • Accuracy
  • Precision / Recall
  • Relevance
  • Faithfulness
  • Groundedness
  • Context relevance
  • Hallucination detection
  • Toxicity and safety
  • Robustness
  • Latency
  • Token consumption
  • Cost per request
RAG & Vector Technologies
  • Experience designing and testing RAG pipelines.
  • Knowledge of vector databases such as:
  • o Azure AI Search
  • o Pinecone
  • o Weaviate
  • o Milvus
  • o Qdrant
  • o FAISS
  • o Elasticsearch/OpenSearch
  • Experience evaluating retrieval quality, embeddings, reranking, chunking, and search strategies.
Cloud & DevOps
Experience with one or more cloud platforms:
  • Microsoft Azure
  • AWS
  • Google Cloud Platform
Good understanding of:
  • Docker
  • Kubernetes
  • CI/CD
  • GitHub Actions / Azure DevOps / Jenkins
  • Infrastructure automation
  • Cloud monitoring and logging
  • Secrets and configuration management
MLOps / LLMOps
  • Experience implementing MLOps/LLMOps practices.
  • Model and prompt versioning.
  • Experiment tracking.
  • Dataset management.
  • Automated evaluation pipelines.
  • Continuous monitoring.
  • Model/application performance tracking.
  • Production deployment and rollback strategies.
Preferred Skills
  • Experience building AI testing platforms or internal AI developer tools.
  • Experience with agent evaluation and multi-agent systems.
  • Experience with MCP servers and MCP-based tool integrations.
  • Knowledge of AI security, prompt injection, jailbreak testing, data leakage, and adversarial evaluation.
  • Experience with synthetic data generation and automated test-case generation.
  • Knowledge of responsible AI and AI governance.
  • Experience with SQL and NoSQL databases.
  • Familiarity with Kafka or event-driven architectures.
  • Experience with TypeScript/Node.js is an advantage.
  • Experience working in Agile/Scrum environments.
Key Deliverables

The Senior AI Harness Engineer will be expected to deliver:

  • Reusable AI testing and evaluation frameworks.
  • Automated LLM and agent regression test suites.
  • AI quality and performance dashboards.
  • Automated evaluation pipelines integrated with CI/CD.
  • Benchmark and golden datasets.
  • RAG and agent evaluation capabilities.
  • Prompt/model comparison frameworks.
  • AI observability and monitoring solutions.
  • Production-ready engineering standards for GenAI applications.
Experience
  • 8+ years of overall software/engineering experience.
  • 4+ years of experience in AI/ML or Generative AI engineering.
  • Strong hands-on experience with Python and AI/LLM technologies.
  • Proven experience developing AI evaluation, testing, automation, or LLMOps solutions.
  • Experience working with enterprise-scale AI applications and production deployments.
Soft Skills
  • Strong analytical and problem-solving skills.
  • Ability to translate AI quality requirements into measurable engineering solutions.
  • Strong communication and collaboration skills.
  • Ability to work with AI researchers, ML engineers, software developers, DevOps engineers, and business stakeholders.
  • Strong ownership mindset with a focus on building reliable and maintainable AI systems.
  • Ability to mentor engineers and provide technical leadership.
Education

Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Engineering, or a related field.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. AI/ML Engineer
Sr. AI/ML Engineer

GECO Asia Pte Ltd • Singapore

Hybrid
SGD 150,000 - 210,000
Applied AI Engineer
Applied AI Engineer

Wilmar International • Singapore

On-site
SGD 120,000 - 160,000
AI Engineer
AI Engineer

ASCENDION ENGINEERING SOLUTIONS SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 180,000
Senior AI Engineer (Principal-Level Scope)
Senior AI Engineer (Principal-Level Scope)

CHEMT BIOTECHNOLOGY PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Full Stack AI Engineer
Full Stack AI Engineer

Amplify Health • Singapore

On-site
SGD 120,000 - 180,000
#EG AI Engineer
#EG AI Engineer

NCS Group • Singapore

On-site
SGD 80,000 - 120,000
AI Developer
AI Developer

OPUS IT SERVICES PTE LTD • Singapore

On-site
SGD 70,000 - 120,000
AI Engineer (Managed Services)
AI Engineer (Managed Services)

Avepoint • Singapore

On-site
SGD 90,000 - 120,000
Flexible working hours
Access to high-performance GPU resources
Continued learning and development opportunities
AI Engineer
AI Engineer

Accenture Southeast Asia • Singapore

On-site
SGD 180,000 - 240,000
AI Engineer (GenAI / RAG / LangGraph / Python)
AI Engineer (GenAI / RAG / LangGraph / Python)

Unison Group • Singapore

On-site
SGD 120,000 - 180,000