Senior/Principal LLM & Evals Engineer

Perennial Resources International

Toronto

Hybrid

CAD 120,000 - 190,000

Full time

11 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Perennial Resources International seeks an experienced GenAI engineer to design and improve reliable AI applications that answer complex financial questions using structured and unstructured data.

The role focuses on LLM evaluation and production-grade systems with a deep understanding of front-office workflows and the standards of accuracy, reliability, and responsiveness expected by investment professionals.

Qualifications

  • Hands-on experience building and operating LLM/GenAI applications in production.
  • Deep knowledge of LLM evaluation methods, benchmarks, test-set design, error analysis, and continuous quality measurement.
  • Strong understanding of front-office financial workflows and accuracy requirements.

Responsibilities

  • Design evaluation frameworks for LLM-powered financial apps focusing on accuracy and grounding.
  • Create datasets and golden sets including difficult and business-critical test cases.
  • Build automated evaluation pipelines with deterministic tests, model evaluation, and financial validation.
  • Reduce hallucinations, unsupported conclusions, and retrieval failures.
  • Develop production-grade applications using structured and unstructured financial data.

Skills

LLM evaluation
Production AI engineering
Financial-domain fluency
Retrieval-augmented generation
Software engineering

Tools

MCP servers
Embeddings

Job description

We are seeking an experienced GenAI engineer with deep expertise in LLM evaluation, production-grade AI systems, and financial services. The successful candidate will design and improve reliable AI applications that answer complex financial questions using structured and unstructured data.

This is primarily an LLM engineering role—not a trading or quantitative research position. However, candidates must understand front-office financial workflows and the standards of accuracy, reliability, and responsiveness expected by investment professionals.

Ideal candidates may come from financial-data and AI platforms, banks, asset managers, fintech companies, or specialist firms such as Rogo, RavenPack, LSEG, FactSet, Bloomberg, S&P Global, or similar organizations.

Key responsibilities
  • Design comprehensive evaluation frameworks for LLM-powered financial applications, covering factual accuracy, relevance, completeness, grounding, consistency, and appropriate refusal behavior.
  • Build representative evaluation datasets and financial “golden sets,” including difficult, ambiguous, adversarial, and business-critical test cases.
  • Establish automated evaluation pipelines combining deterministic tests, model-based evaluation, human review, and expert financial validation.
  • Identify and reduce hallucinations, unsupported conclusions, citation errors, retrieval failures, and other sources of unreliable output.
  • Develop production-grade applications that answer questions using structured and unstructured financial information, including research, filings, transcripts, news, market data, and proprietary documents.
  • Design and optimize retrieval, RAG, search, ranking, reranking, context construction, and tool-use workflows.
  • Build or integrate MCP servers and other interfaces that allow models to query enterprise data, research systems, analytical tools, and external services securely.
  • Evaluate the complete application rather than the model in isolation, including retrieval quality, tool selection, tool execution, context quality, answer generation, citations, and end‑to‑end task success.
  • Implement monitoring and observability for production AI systems, including quality degradation, data drift, latency, cost, failure modes, and user feedback.
  • Apply fine‑tuning, supervised learning, preference optimization, synthetic‑data generation, and related techniques where they provide measurable improvements.
  • Develop bespoke smaller language models through model distillation—for example, transferring capabilities from a large teacher model to a smaller model optimized for a particular financial domain or workflow.
  • Make informed decisions about when to use prompting, retrieval, tools, fine‑tuning, distillation, or a combination of techniques.
  • Optimize models and applications for accuracy, latency, scalability, inference cost, security, and operational reliability.
  • Work closely with financial‑domain experts, product teams, data engineers, and application engineers to translate front‑office requirements into measurable AI capabilities.
  • Define production‑readiness criteria, release gates, regression tests, and quality thresholds for new models and application changes.
Required experience
  • Significant hands‑on experience building and operating LLM or GenAI applications in production.
  • Deep knowledge of LLM evaluation methods, benchmarks, test‑set design, error analysis, regression testing, and continuous quality measurement.
  • Demonstrated ability to improve the factual accuracy and reliability of answers generated from enterprise or domain‑specific information.
  • Strong experience working with unstructured information and document‑heavy workflows.
  • Practical expertise in retrieval‑augmented generation, semantic and hybrid search, embeddings, reranking, grounding, citations, and context management.
  • Experience building tool‑using or agentic systems, ideally including MCP servers or comparable model‑to‑data and model‑to‑tool interfaces.
  • Hands‑on experience with model fine‑tuning, distillation, or the development of domain‑specific small and large language models.
  • Strong understanding of the trade‑offs among model quality, model size, latency, throughput, infrastructure requirements, and inference cost.
  • Experience implementing production monitoring, evaluation pipelines, quality controls, and safeguards for probabilistic AI systems.
  • Meaningful financial‑services experience, with a working understanding of front‑office users, terminology, data, and workflows.
  • Familiarity with how investment professionals consume research, interrogate financial information, evaluate evidence, and make time‑sensitive decisions.
  • Strong software‑engineering skills and the ability to build robust, testable, maintainable production systems.
Candidates should understand at least several of the following:
  • Equity, fixed‑income, credit, commodities, foreign‑exchange, or multi‑asset workflows.
  • Financial research, company analysis, market intelligence, news analytics, and investment decision support.
  • Company filings, earnings transcripts, financial statements, estimates, corporate actions, and market data.
  • The distinction between facts, estimates, opinions, forecasts, and model‑generated conclusions.
  • Data entitlements, auditability, source attribution, information security, and regulatory or compliance considerations.
  • The importance of timeliness, point‑in‑time correctness, provenance, and reproducibility in financial applications.

Direct experience as a trader or quant is not required. The essential requirement is sufficient domain fluency to understand front‑office use cases and recognize when an AI‑generated answer is incomplete, misleading, unsupported, or financially implausible.

Core competencies
  • LLM evaluation: Can define what “good” means, measure it rigorously, and create repeatable systems for improving quality.
  • Production AI engineering: Can take an LLM application from prototype to a reliable, observable, scalable production service.
  • Financial‑domain fluency: Understands front‑office users, financial information, relevant workflows, and the consequences of inaccurate output.
  • Grounded answer generation: Knows how to produce accurate, traceable responses supported by authoritative sources.
  • Retrieval and data integration: Can connect models effectively to structured data, unstructured content, search systems, tools, and enterprise platforms.
  • Model optimization: Understands fine‑tuning, distillation, model selection, and the construction of specialized models for defined tasks.
  • Analytical problem‑solving: Can diagnose whether failures originate in the underlying data, retrieval, tool use, context construction, prompting, model behavior, or application logic.
  • Risk and quality mindset: Anticipates edge cases and designs controls appropriate for high‑stakes financial applications.
  • Cross‑functional collaboration: Communicates effectively with financial experts, engineers, researchers, product leaders, and senior stakeholders.
  • Outcome orientation: Focuses on measurable user and business outcomes rather than model demonstrations alone.
Preferred experience
  • Experience at a financial‑data provider, investment‑research platform, financial AI company, bank, asset manager, hedge fund, or capital‑markets technology business.
  • Experience supporting research analysts, portfolio managers, traders, salespeople, or other front‑office professionals.
  • Experience building domain‑specific models or AI applications for financial research and decision support.
  • Experience with SaaS products, multi‑tenant platforms, enterprise deployments, or customer‑facing AI applications.
  • Familiarity with model governance, access controls, data privacy, audit trails, and regulated production environments.
  • Experience evaluating both proprietary and open‑weight models and selecting the appropriate model for a given task.
What success looks like

The successful candidate will build AI systems that financial professionals can use with confidence. Their work will result in measurable improvements in answer accuracy, grounding, coverage, latency, cost, and reliability. They will also establish a disciplined evaluation process that makes quality visible, catches regressions before release, and supports the development of specialized models for high‑value financial workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Engineer - SRE, LLM Agents, Full-Stack Architecture
Lead AI Engineer - SRE, LLM Agents, Full-Stack Architecture

iVedha Inc. • Canada

On-site
CAD 120,000 - 150,000
Opportunity for architectural influence
Working with cutting-edge AI models
Leadership opportunities
Copy of AI Engineer
Copy of AI Engineer

Jobless • Toronto

On-site
CAD 90,000 - 130,000
Gen AI Engineering and Scaled AI Transformation
Gen AI Engineering and Scaled AI Transformation

Citi • Mississauga

On-site
CAD 120,000 - 160,000
Applied AI Engineer
Applied AI Engineer

Compunnel, Inc. • Montreal (administrative region)

On-site
CAD 100,000 - 140,000
AI Engineer
AI Engineer

Fulfillment IQ • Toronto

On-site
CAD 135,000 - 170,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Hallmark Global Solutions Ltd • Mississauga

On-site
CAD 100,000 - 150,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank | Canada's Challenger Bank • Toronto

On-site
CAD 150,000 - 210,000
AI Engineer - Assistant Vice President
AI Engineer - Assistant Vice President

Citigroup Inc. • Mississauga

On-site
CAD 131,000 - 198,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

Kinvie • Toronto

On-site
CAD 140,000 - 190,000