AI Engineer Sr

Dayforce

Toronto

Hybrid

CAD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Dayforce is seeking an AI Engineer - Agentic Systems Evaluation to help define how we measure, test, and improve the quality of enterprise AI systems. You will design evaluation frameworks for agentic AI, multi-agent workflows, and RAG-driven systems, tackling reliability, cost, and safety.

Join a team building scalable evaluation infrastructure and producing rigorous benchmarks to guide production deployments.

Qualifications

  • Experience with Python and software engineering.
  • LLM application development experience.
  • AI/ML evaluation and experimentation.
  • Automated testing and data analysis experience.
  • Familiarity with RAG, embeddings, hybrid search, and reranking.
  • Experience with tool/function calling and structured outputs.
  • Agent orchestration and multi-agent workflows.
  • AI observability and tracing.
  • APIs and enterprise integrations.

Responsibilities

  • Build Agentic Evaluation Frameworks covering the full AI execution lifecycle.
  • Design experiments and benchmarks for architecture comparisons (single-agent vs multi-agent).
  • Develop automated evaluation and regression testing infrastructure.
  • Analyze agent traces, tool usage, and failure modes to improve reliability and cost.

Skills

Python
LLM development
AI evaluation
Automated testing
RAG and embeddings
Tool calls and structured outputs
Agent orchestration
AI observability
APIs and enterprise integrations

Education

Bachelor's degree in CS/Engineering

Tools

Agents SDK
LangGraph/LangChain

Job description

About the opportunity

We are looking for an AI Engineer - Agentic Systems Evaluation to help define how we measure, test, and improve the quality of enterprise AI systems.

As AI evolves from conversational assistants and Retrieval-Augmented Generation (RAG) into tool using agents, multi-agent systems, and autonomous business workflows, evaluating only the final response is no longer enough.

An agent may reach the right answer while choosing the wrong tool, taking unnecessary steps, retrieving incorrect context, failing to escape to a human or violating business rules.

This role will build the evaluation frameworks needed to understand not only whether an AI system succeeded, but how it succeeded, how reliably it can repeat that outcome, and whether the architecture is appropriate for the problem.

What you’ll get to do

Build Agentic Evaluation Frameworks

Design evaluation methodologies covering the complete AI execution lifecycle:

Intent Planning Retrieval Tool Use Reasoning Action Business Outcome

Evaluate systems across dimensions including:

  • task and business outcome accuracy
  • planning and decision quality
  • retrieval quality and groundedness
  • tool selection and execution
  • agent routing and delegation
  • human escalation decisions
  • reliability and failure recovery
  • safety and policy compliance
  • latency, token consumption, and cost

Evaluate Agentic Design Patterns

Design experiments and benchmarks that help engineering teams determine which architecture works best for a given problem.

Evaluate patterns such as:

  • Single-agent vs. multi-agent systems
  • Supervisor/router architectures
  • Planner–executor patterns
  • Sequential and parallel workflows
  • Tool-using agents
  • Human-in-the-loop workflows
  • Long-running agents with state and memory

Measure whether additional agent complexity actually improves task success, reliability, and business outcomes enough to justify increased latency, cost, and operational complexity.

Advance RAG & Knowledge Evaluation

Build rigorous evaluation approaches for enterprise retrieval and knowledge systems, including:

  • retrieval precision and relevance
  • groundedness and citation accuracy
  • source authority and freshness
  • chunking, metadata, and indexing strategies
  • semantic vs. hybrid search and reranking
  • permission-aware retrieval

Evaluate how retrieval decisions ultimately impact downstream agent performance, rather than treating RAG evaluation as an isolated problem.

Build Automated Evaluation & Regression Testing

Develop scalable evaluation infrastructure including:

  • golden and synthetic datasets
  • scenario and adversarial test suites
  • deterministic graders
  • LLM-as-a-Judge evaluation
  • human evaluation workflows
  • trace-based evaluation
  • automated regression testing

Integrate evaluations into AI development and release pipelines so changes to models, prompts, retrieval, tools, or agent architectures can be measured before reaching production.

Build Agent Trace & Failure Analysis

Analyze complete agent execution traces including planning, retrieved context, tool calls, handoffs, retries, exceptions, latency, and cost.

Develop failure taxonomies that distinguish between:

Model | Retrieval | Planning | Tool | Routing | Memory | Integration | Policy | Orchestration failures

Turn production failures and user feedback into measurable regression tests and engineering improvements.

Define Production AI Quality

Establish measurable quality standards and release criteria for AI systems.

Metrics may include Task Success Rate, First-Pass Success Rate, Tool Selection Accuracy, Agent Routing Accuracy, Plan Execution Fidelity, Failure Recovery Rate, Human Escalation Accuracy, Business Outcome Accuracy, Cost / Latency per Successful Task

Help teams answer a fundamental question:

Is this AI system reliable, safe, efficient, and valuable enough to operate in production?

Skills and experience we value

Strong experience with:

  • Python and software engineering
  • LLM application development
  • AI/ML evaluation and experimentation
  • automated testing and data analysis
  • RAG, embeddings, hybrid search, and reranking
  • tool/function calling and structured outputs
  • agent orchestration and multi-agent workflows
  • AI observability and tracing
  • APIs and enterprise integrations

Experience with platforms or frameworks such as Agents SDK, LangGraph/LangChain, Microsoft AI Foundry, Amazon Bedrock, or similar agent platforms is valuable.

You should be comfortable working with evaluation techniques such as offline/online evals, deterministic graders, model-based graders, human evaluation, synthetic datasets, adversarial testing.

What would make you stand out

You don't stop when an agent successfully completes a task. You ask:

  • Did it choose the right approach?
  • Did it use the right tools and information?
  • Can it succeed consistently?
  • Can it recover when something fails?
  • Did adding more agents improve the outcome?
  • Could a simpler architecture achieve the same result?
  • What did the successful outcome cost?

You turn those questions into measurable experiments that help engineering teams build better AI systems.

Why This Role Matters

Enterprise AI is moving from systems that answer questions to systems that make decisions and perform work.

That changes how quality must be measured.

The next generation of AI evaluation must measure:

What the system understood what it retrieved what it decided what it did and whether the business outcome was correct.

This role will help establish the engineering discipline required to make those systems measurable, reliable, and production-ready.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Agentic AI Optimization Developer
Agentic AI Optimization Developer

Equifax, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior AI Evaluation Engineer - Agentic Systems
Senior AI Evaluation Engineer - Agentic Systems

Dayforce • Toronto

Hybrid
CAD 120,000 - 180,000
AI Quality Engineer
AI Quality Engineer

Centraprise • Vancouver

On-site
CAD 80,000 - 100,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Kinvie • Toronto

On-site
CAD 90,000 - 120,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

EQ Bank | Canada's Challenger Bank • Toronto

On-site
CAD 120,000 - 180,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
Senior AI Engineer – JVM Engineering (Java/Kotlin), Knowledge Graphs, MCP
Senior AI Engineer – JVM Engineering (Java/Kotlin), Knowledge Graphs, MCP

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Toronto

Hybrid
CAD 120,000 - 170,000
Principal AI Architect
Principal AI Architect

Harnham • Toronto

Hybrid
CAD 200,000 - 220,000
Flexible hybrid working environment
Competitive CAD 200k–220k salary
Forward Deployed AI Engineer
Forward Deployed AI Engineer

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
Software Engineer – AI
Software Engineer – AI

Webhosting • Quebec

On-site
CAD 90,000 - 140,000