Agentic AI Optimization Developer

United States Digital Space LLC

Toronto

On-site

CAD 109,000 - 143,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is building autonomous AI agents for enterprise-grade tasks, focusing on safety, reliability, and cost-effective operation. You will be responsible for evaluating agent behavior, tuning prompts, and ensuring robust guardrails in production.

The role emphasizes automated evaluation, tool-calling optimization, and collaboration with AI leads. It targets engineers with strong ML QA backgrounds and Python scripting abilities.

Qualifications

  • 3+ years in software quality engineering, test automation, or data/ML engineering focused on LLM testing.
  • Hands-on with LLM orchestration frameworks (LangGraph, ADKs or internal SDKs).
  • Deep JSON schema design for LLM tool- and function-calls, and structured outputs.
  • Automated testing and tracing/observability for asynchronous, non-deterministic systems.

Responsibilities

  • Build golden datasets (documents, queries, agent trajectories) for testing.
  • Design and run automated evaluation pipelines for accuracy, latency, and token spend.
  • Audit agent trajectories to identify deviations in decision paths.
  • Refine prompts, contexts, and few-shot examples to optimize behavior.
  • Tune tool calls and external APIs to minimize errors and token overhead.
  • Collaborate with AI leads and developers to feedback evaluation insights.

Skills

LLM orchestration
JSON schema design
Python scripting
Automated testing
Observability

Tools

LangGraph
ADKs
UiPath Maestro
Vertex AI
Gemini Enterprise

Job description

  • Toronto
  • Canada
  • Technology
  • Full time
  • 8/12/2026
  • J00178408

To play this video, please accept Functional \\& Personalisation cookies in your cookie preferences.


Manage Preferences


the company is where you can power your possible. If you want to achieve your true potential, chart new paths, develop new skills, collaborate with bright minds, and make a meaningful impact, we want to hear from you.


Synopsis of the role

At the company, we are moving past passive AI chat interfaces to build the future of autonomous workflows. We are creating intelligent, self-correcting multi-agent systems that can navigate complex software environments, utilize external tools, and solve open-ended business problems with minimal human intervention. To ensure these systems are safe, reliable, and enterprise-grade, we are seeking an analytical Agentic AI Evaluation \\& Tuning Engineer.


In this role, you will be the guardian of our production AI reliability. You will bridge the gap between raw Large Language Model (LLM) capabilities and flawless autonomous execution. Unlike traditional software testers or prompt engineers, you will focus on the behavior, decision‑making logic, tool‑use efficiency, and long‑term stability of multi‑agent architectures. Your mission is to build the automated evaluation frameworks that keep our agents accurate, cost‑effective, and hallucination‑free.


What you will do

Golden Dataset Curation \\& Automated Evaluation



  • Build the \"Golden Set\": Curate, maintain, and augment high‑quality reference datasets (Golden Sets) of documents, user queries, and expected agent trajectories to serve as the ultimate source of truth for testing.

  • Automate Eval Cycles: Design and implement automated, continuous evaluation pipelines to measure agent accuracy, latency, token spend, and fallback reliability before code hits production.

  • Trajectory \\& Reasoning Auditing: Trace and dissect complex, multi‑step agent \"thought\" processes (e.g., ReAct, Reflection loops) to pinpoint exactly where an agent deviates from its intended logic path.


Agent Tuning \\& Developer Collaboration



  • Behavioral Optimization: Refine system prompts, context windows, and few‑shot examples to optimize how agents execute complex, multi‑step workflows.

  • Tool \\& Function-Calling Optimization: Fine‑tune how agents interact with external APIs, databases, and UiPath RPA workflows—minimizing execution errors, redundant calls, and token overhead.

  • Augment Development: Partner closely with AI Solution Leads and AI Agent Developers to feed evaluation insights back into the development lifecycle, helping them build robust, reusable, and self‑correcting agent components.


Production Guardrails \\& Lifecycle Management (LLMOps)



  • Defeat Drift \\& Hallucinations: Actively monitor deployed agents to identify, troubleshoot, and mitigate semantic drift, prompt injections, infinite execution loops, and hallucinations.

  • Maintain Autonomous Integrity: Implement robust guardrail frameworks to ensure agents maintain reliable, fact‑based autonomous decision‑making post‑deployment in production.

  • RAG \\& Knowledge Integration: Optimize Domain‑Specific Knowledge Bases and Retrieval-Augmented Generation (RAG) pipelines to ensure agents pull from accurate data rather than assumptions.


What Experience You Need


  • Experience: 3+ years of professional experience in software quality engineering, test automation, or data/ML engineering, with a dedicated focus on LLM testing, prompt tuning, or orchestration patterns over the last 1–2 years.

  • Agentic \\& LLM Frameworks: Proven hands‑on experience working with LLM orchestration frameworks (e.g., LangGraph, ADKs or specialized internal SDKs).

  • Function Calling Mastery: Deep understanding of JSON schema design for LLM tool‑calling, function‑calling, and structured outputs.

  • Advanced Debugging \\& Automation: Strong background in writing automated test scripts (Python‑heavy) and using tracing/observability concepts to debug cascading errors in asynchronous, non‑deterministic systems.


What Could Set You Apart


  • Experience with AI evaluation and observability platforms

  • Live production experience testing Agentic workflows and GenAI solutions

  • Familiarity with Google Cloud AI suite (Vertex \\& Gemini Enterprise Agent Platform) and UiPath ecosystem (Maestro).

  • Experience utilizing LLMs to securely generate high‑quality synthetic data for edge‑case testing.

  • Proficiency in Python or TypeScript, with a deep understanding of asynchronous programming, API design, and microservices architecture.

  • Demonstrated learning agility and a proactive approach to mastering new technologies.


This is a newly created position.


This role has a base salary range of $78,504 - $103,037 The final salary selected from this range is dependent on the location of the role as well as the experiences \\& capabilities of the candidate selected. We offer comprehensive compensation and healthcare packages, paid time off, and organizational growth potential through our online learning platform with guided career tracks.


Are you ready to power your possible? At the company, we value and celebrate diversity. We are committed to fostering an inclusive, equitable, and accessible workplace where every member of our team feels respected, supported, and has the opportunity to reach their full potential. We strongly encourage applications from people with disabilities. Accommodations are available upon request for candidates participating in all aspects of the recruitment process. For a confidential request, please contact your recruiter or email us at hr@unitedstatesdigital.space to make arrangements. If you have any questions regarding accessibility at the company, you can also contact us at this same email address, hr@unitedstatesdigital.space.


Who is the company?

At the company, we believe knowledge drives progress. As a global data, analytics and technology company, we play an essential role in the global economy by helping employers, employees, financial institutions and government agencies make critical decisions with greater confidence.


We work to help create seamless and positive experiences during life’s pivotal moments: applying for jobs or a mortgage, financing an education or buying a car. Our impact is real and to accomplish our goals we focus on nurturing our people for career advancement and their learning and development, supporting our next generation of leaders, maintaining an inclusive and diverse work environment, and regularly engaging and recognizing our employees. Regardless of location or role, the individual and collective work of our employees makes a difference and we are looking for talented team players to join us as we help people live their financial best.


the company is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agentic AI Solutions Architect - Vice President
Agentic AI Solutions Architect - Vice President

Citi • Mississauga

On-site
CAD 120,000 - 171,000
Sr. MTS Software Engineering - AIRI
Sr. MTS Software Engineering - AIRI

United States Digital Space LLC • Toronto

On-site
CAD 140,000 - 210,000
Forward Deployed Agile Software Engineer
Forward Deployed Agile Software Engineer

United States Digital Space LLC • Toronto

On-site
CAD 120,000 - 180,000
Senior Lead Systems Engineer, AI & Automation
Senior Lead Systems Engineer, AI & Automation

United States Digital Space LLC • Toronto

Hybrid
CAD 119,000 - 187,000
Equity grants
Hybrid work model
Comprehensive benefits
+2
Principal Engineer, Agentic Products & Workflows
Principal Engineer, Agentic Products & Workflows

United States Digital Space LLC • Toronto

On-site
CAD 175,000 - 200,000
Group retirement savings plan matching
Health, dental, vision benefits
Wellness spending account
+3
AI-Augmented Full-Stack Engineer - Vice President
AI-Augmented Full-Stack Engineer - Vice President

United States Digital Space LLC • Mississauga

On-site
CAD 168,000 - 237,000
Agentic AI Solutions Architect - Vice President
Agentic AI Solutions Architect - Vice President

Citigroup Inc. • Mississauga

On-site
CAD 120,000 - 171,000
Senior Software Engineer, Applied AI (Toronto)
Senior Software Engineer, Applied AI (Toronto)

United States Digital Space LLC • Toronto

Hybrid
CAD 188,000 - 242,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Citi • Mississauga

On-site
CAD 168,000 - 237,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Citibank (Switzerland) AG • Mississauga

Hybrid
Confidential