Agentic AI Optimization Developer

Equifax, Inc.

Toronto

On-site

CAD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Equifax, Inc. is seeking an analytical Agentic AI Evaluation & Tuning Engineer to safeguard production AI reliability and autonomous workflows.

You will bridge raw LLM capabilities with flawless autonomous execution, focusing on behavior, decision-making, and tool usage efficiency. You will build automated evaluation frameworks to measure accuracy, latency, and token spend while auditing reasoning paths and guarding against hallucinations.

Qualifications

  • 3+ years in software quality engineering or data/ML engineering focused on LLM testing.
  • Hands-on with LLM orchestration frameworks (LangGraph, ADKs or internal SDKs).
  • Deep understanding of JSON schema design for LLM tool‑calling and structured outputs.
  • Strong Python scripting and tracing/observability for async systems.

Responsibilities

  • Curate Golden Sets and maintain high-quality datasets for testing.
  • Design automated evaluation pipelines measuring accuracy, latency, token spend, and reliability.
  • Trace multi-step agent reasoning to pinpoint deviations in logic path.
  • Tune prompts, context windows, and tool-calling to improve robustness; collaborate with AI teams.

Skills

LLM testing
Test automation
Python scripting
Observability

Tools

LangGraph
ADKs
SDKs
JSON schema

Job description

Synopsis of the role

At Equifax, we are moving past passive AI chat interfaces to build the future of autonomous workflows. We are creating intelligent, self‑correcting multi‑agent systems that can navigate complex software environments, utilize external tools, and solve open-ended business problems with minimal human intervention. To ensure these systems are safe, reliable, and enterprise‑grade, we are seeking an analytical Agentic AI Evaluation & Tuning Engineer.

In this role, you will be the guardian of our production AI reliability. You will bridge the gap between raw Large Language Model (LLM) capabilities and flawless autonomous execution. Unlike traditional software testers or prompt engineers, you will focus on the behavior, decision‑making logic, tool‑use efficiency, and long‑term stability of multi‑agent architectures. Your mission is to build the automated evaluation frameworks that keep our agents accurate, cost‑effective, and hallucination‑free.

What you will do
Golden Dataset Curation & Automated Evaluation
  • Build the "Golden Set": Curate, maintain, and augment high‑quality reference datasets (Golden Sets) of documents, user queries, and expected agent trajectories to serve as the ultimate source of truth for testing.

  • Automate Eval Cycles: Design and implement automated, continuous evaluation pipelines to measure agent accuracy, latency, token spend, and fallback reliability before code hits production.

  • Trajectory & Reasoning Auditing: Trace and dissect complex, multi‑step agent "thought" processes (e.g., ReAct, Reflection loops) to pinpoint exactly where an agent deviates from its intended logic path.

Agent Tuning & Developer Collaboration
  • Behavioral Optimization: Refine system prompts, context windows, and few‑shot examples to optimize how agents execute complex, multi‑step workflows.

  • Tool & Function‑Calling Optimization: Fine‑tune how agents interact with external APIs, databases, and UiPath RPA workflows—minimizing execution errors, redundant calls, and token overhead.

  • Augment Development: Partner closely with AI Solution Leads and AI Agent Developers to feed evaluation insights back into the development lifecycle, helping them build robust, reusable, and self‑correcting agent components.

Production Guardrails & Lifecycle Management (LLMOps)
  • Defeat Drift & Hallucinations: Actively monitor deployed agents to identify, troubleshoot, and mitigate semantic drift, prompt injections, infinite execution loops, and hallucinations.

  • Maintain Autonomous Integrity: Implement robust guardrail frameworks to ensure agents maintain reliable, fact‑based autonomous decision‑making post‑deployment in production.

  • RAG & Knowledge Integration: Optimize Domain‑Specific Knowledge Bases and Retrieval‑Augmented Generation (RAG) pipelines to ensure agents pull from accurate data rather than assumptions.

What Experience You Need
  • Experience: 3+ years of professional experience in software quality engineering, test automation, or data/ML engineering, with a dedicated focus on LLM testing, prompt tuning, or orchestration patterns over the last 1–2 years.

  • Agentic & LLM Frameworks: Proven hands‑on experience working with LLM orchestration frameworks (e.g., LangGraph, ADKs or specialized internal SDKs).

  • Function Calling Mastery: Deep understanding of JSON schema design for LLM tool‑calling, function‑calling, and structured outputs.

  • Advanced Debugging & Automation: Strong background in writing automated test scripts (Python‑heavy) and using tracing/observability concepts to debug cascading errors in asynchronous, non‑deterministic systems.

What Could Set You Apart
  • Experience with AI evaluation and observability platforms

  • Live production experience testing Agentic workflows and GenAI solutions

  • Familiarity with Google Cloud AI suite (Vertex & Gemini Enterprise Agent Platform) and UiPath ecosystem (Maestro).

  • Experience utilizing LLMs to securely generate high‑quality synthetic data for edge‑case testing.

  • Proficiency in Python or TypeScript, with a deep understanding of asynchronous programming, API design, and microservices architecture.

  • Demonstrated learning agility and a proactive approach to mastering new technologies.

This is a newly created position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agentic AI Optimization Developer
Agentic AI Optimization Developer

United States Digital Space LLC • Toronto

On-site
CAD 109,000 - 143,000
Sr. Software Development Engineer in Test (Agentic)
Sr. Software Development Engineer in Test (Agentic)

United States Digital Space LLC • Vancouver

On-site
CAD 110,000 - 150,000
Copy of AI Engineer
Copy of AI Engineer

Jobless • Toronto

On-site
CAD 90,000 - 130,000
Agentic AI Solutions Architect - Vice President
Agentic AI Solutions Architect - Vice President

Citi • Mississauga

On-site
CAD 120,000 - 171,000
AI Evaluation Engineer
AI Evaluation Engineer

United States Digital Space LLC • Kitchener

On-site
CAD 90,000 - 130,000
Automation Solution Lead
Automation Solution Lead

Equifax, Inc. • Toronto

On-site
CAD 120,000 - 180,000
AI Engineer
AI Engineer

Valsoft Corporation • Canada

On-site
CAD 100,000 - 140,000
Agentic AI Solutions Architect - Vice President
Agentic AI Solutions Architect - Vice President

Citigroup Inc. • Mississauga

On-site
CAD 120,000 - 171,000
AI Quality Engineer
AI Quality Engineer

Centraprise • Vancouver

On-site
CAD 80,000 - 100,000
AI Engineer - Canada
AI Engineer - Canada

Pulsora • Canada

Remote
CAD 60,000 - 70,000