Avensys is a reputed global IT professional services company headquartered in Singapore. Our service spectrum includes enterprise solution consulting, business intelligence, business process automation and managed services. Given our decade of success we have evolved to become one of the top trusted providers in Singapore and service a client base across banking and financial services, insurance, information technology, healthcare, retail, and supply chain.
We are currently looking to hire Senior AI Harness Engineer. This is an exciting opportunity to expand your skill set, achieve job satisfaction and work-life balance. More details as below.
Role Overview
We are looking for a Lead AI Harness Engineer to lead the design, development, and implementation of AI/LLM evaluation and orchestration frameworks. The role involves building robust AI harnesses to test, evaluate, monitor, and improve AI agents, LLM applications, and agentic workflows in production environments.
Key Responsibilities
- Lead the design and development of AI/LLM harness frameworks for testing, evaluation, validation, and monitoring.
- Develop automated frameworks to evaluate LLM responses, AI agents, RAG pipelines, and agentic workflows.
- Design test scenarios covering accuracy, relevance, hallucination, safety, latency, reliability, and performance.
- Build evaluation pipelines using Python and modern AI/ML frameworks.
- Develop and integrate LLM-as-a-Judge, benchmark datasets, evaluation metrics, and automated scoring mechanisms.
- Create test harnesses for multi-agent systems, tool calling, function calling, and autonomous workflows.
- Work with Generative AI, RAG, embeddings, vector databases, and prompt engineering.
- Integrate AI harnesses with CI/CD pipelines to enable continuous AI testing and regression testing.
- Establish best practices for AI quality, observability, governance, and responsible AI.
- Analyze evaluation results and identify areas for improving model prompts, workflows, retrieval strategies, and agent behavior.
- Lead technical discussions and mentor engineers working on AI/ML and GenAI solutions.
- Collaborate with product managers, data scientists, AI engineers, DevOps, and business stakeholders.
- Evaluate emerging AI technologies, models, frameworks, and tooling and recommend suitable solutions.
Required Technical Skills
- Strong programming experience in Python.
- Strong hands‑on experience with Generative AI, LLMs, and AI Agents.
- Experience building AI/LLM evaluation or testing harnesses.
- Strong knowledge of RAG, prompt engineering, embeddings, vector databases, and LLM orchestration.
- Experience with frameworks such as LangChain, LangGraph, LlamaIndex, Semantic Kernel, or equivalent.
- Experience with LLM evaluation frameworks such as Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, or equivalent.
- Experience with LLM‑as‑a‑Judge and automated evaluation methodologies.
- Knowledge of REST APIs, microservices, JSON, and tool/function calling.
- Experience with Docker, Kubernetes, CI/CD, Git, and cloud platforms.
- Familiarity with Azure AI, AWS Bedrock, Google Vertex AI, Azure OpenAI, or equivalent AI platforms.
- Understanding of MLOps/LLMOps, observability, model monitoring, and AI governance.
- Leadership Skills
- Proven experience leading AI/GenAI engineering teams or technical initiatives.
- Ability to define technical architecture and establish engineering standards.
- Strong problem‑solving and analytical skills.
- Ability to translate business requirements into scalable AI solutions.
- Excellent communication and stakeholder‑management skills.
- Experience mentoring engineers and conducting technical reviews.
Preferred Skills
- Experience with Agentic AI and multi‑agent architectures.
- Knowledge of MCP (Model Context Protocol) and AI tool integrations.
- Experience with Azure AI Search, Pinecone, Weaviate, Milvus, Elasticsearch, or similar vector databases.
- Knowledge of AI security, prompt injection, jailbreak testing, and responsible AI.
- Experience implementing AI red‑teaming and adversarial testing.
- Familiarity with TypeScript/Node.js and modern AI application development.
- Experience working with enterprise‑scale AI platforms and production deployments.
- Typical Experience
8–12+ years of overall software/AI engineering experience, with 3+ years of hands‑on Generative AI/LLM experience and proven experience leading AI engineering initiatives.
WHAT’S ON OFFER
You will be remunerated with an excellent base salary and entitled to attractive company benefits. Additionally, you will get the opportunity to enjoy a fun and collaborative work environment, alongside a strong career progression.