Staff AI Engineer

Arcadia

Toronto

Hybrid

CAD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid work up to 3 days per week

Job summary

Arcadia in Toronto is seeking a Staff-level software engineer to define the architecture and technical direction for large-scale AI infrastructure across inference, agentic systems, distributed platforms, and cloud-native backend services.

You’ll help build the infrastructure used to deploy and operate text, voice, vision, code, and domain-specific models, while architecting the runtime, orchestration, safety, and developer tooling required for autonomous AI agents to operate reliably in

Qualifications

  • 8+ years of software engineering experience with large-scale backend or distributed systems.
  • 3+ years building production AI systems, LLM apps, or agentic platforms.
  • Ownership of architecture and delivery of complex technical initiatives.
  • Strong Python and/or TypeScript experience.
  • Experience with Kubernetes, Docker, AWS or GCP and cloud-native deployment.
  • Knowledge of inference optimization and security in agentic systems.

Responsibilities

  • Define the technical architecture and direction for production AI platforms across inference and agentic systems.
  • Architect and build runtime, orchestration, and developer tooling for autonomous AI agents.
  • Design multi-agent coordination systems enabling reasoning, collaboration, and tool use.
  • Build multi-model serving infra across text, voice, code, vision, and domain models.
  • Own the complete model lifecycle including deployment, monitoring, and updates.
  • Optimize inference performance: latency, throughput, reliability, cost.
  • Develop secure tool-use infra for safe API/database interactions.
  • Create guardrails for privacy, sandboxing, prompt safety, and human-in-the-loop oversight.
  • Develop observability, evaluation frameworks to monitor agent behavior and diagnose failures.
  • Design scalable backend services, APIs, event-driven systems, and workflows.
  • Develop SDKs, APIs, and platform capabilities for rapid internal deployment.
  • Lead cross-functional technical initiatives with Product, ML, Infrastructure, and Security.
  • Establish engineering standards for agent design, model serving, and reliability.
  • Mentor senior engineers and raise the technical bar across the org.

Skills

Python
TypeScript
Distributed systems
System design
Kubernetes
Docker
Cloud platforms (AWS/GCP)
High-performance inference
Mentoring

Tools

Kubernetes
Docker
AWS
GCP

Job description

Location: Toronto, ON (Hybrid - x3 per week in downtown Toronto office)

Type: Full-Time

We're partnering with a highly technical AI organization building the infrastructure that powers production AI systems at massive scale. Operating as a AI innovation startup within a much larger global technology business, the team combines the pace, ownership, and greenfield engineering opportunities of an early-stage company with the resources and reach of an established platform serving hundreds of millions of users.

This is not an AI research role. We're looking for a Staff-level software engineer who can define the architecture and technical direction for large-scale AI infrastructure across inference, agentic systems, distributed platforms, and cloud-native backend services.

You’ll help build the infrastructure used to deploy and operate text, voice, vision, code, and domain-specific models, while architecting the runtime, orchestration, safety, and developer tooling required for autonomous AI agents to operate reliably in production.

This is a highly hands-on Staff position. You’ll solve complex engineering problems, lead major technical initiatives, establish platform standards, and influence how production AI systems are built across the organization.

What You'll Do
  • Define the technical architecture and direction for production AI platforms across inference and agentic systems
  • Architect and build the runtime, orchestration, and developer tooling required for autonomous AI agents
  • Design multi-agent coordination systems that enable agents to reason, collaborate, use tools, and execute complex workflows
  • Build multi-model serving infrastructure across text, voice, code, vision, and domain-specific models
  • Own the complete model lifecycle, including deployment, serving, monitoring, updating, routing, and model swapping
  • Optimize inference performance across latency, throughput, reliability, and cost using batching, caching, quantization, and intelligent routing
  • Build secure tool-use infrastructure that allows agents to interact safely with APIs, databases, and internal services
  • Develop guardrails covering permissioning, sandboxing, prompt injection, data leakage, and human-in-the-loop oversight
  • Build evaluation, observability, and monitoring frameworks that measure agent behaviour, detect regressions, and diagnose non-deterministic failures
  • Design scalable backend services, APIs, event-driven systems, and durable workflows supporting production AI applications
  • Develop SDKs, APIs, and platform capabilities that allow internal teams to build and deploy AI agents quickly and safely
  • Lead complex, cross-functional technical initiatives in partnership with Product, ML, Infrastructure, and Security teams
  • Establish engineering standards and best practices for agent design, model serving, tool calling, evaluation, and production reliability
  • Mentor senior engineers and raise the technical bar across the broader engineering organization
What We're Looking For
  • 8+ years of software engineering experience, including significant experience building large-scale backend or distributed systems
  • 3+ years of experience building production AI systems, LLM applications, agentic platforms, or machine learning infrastructure
  • Demonstrated experience owning the architecture and delivery of complex, business-critical technical initiatives
  • Strong understanding of LLM-based agent architectures, including tool use, memory, planning, multi-step workflows, and multi-agent coordination
  • Experience building highly reliable distributed systems using event-driven architectures, task queues, state management, and durable workflows
  • Experience evaluating production LLM systems, building automated evaluations, detecting regressions, and debugging non-deterministic failures
  • Strong programming experience with Python and/or TypeScript, with the ability and willingness to work across both
  • Experience with Kubernetes, Docker, AWS or GCP, and modern cloud-native deployment practices
  • Experience working with commercial LLM APIs, open-source models, or model-serving technologies
  • Understanding of inference optimization techniques such as quantization, batching, caching, routing, and GPU utilization
  • Strong understanding of the security risks associated with agentic systems, including prompt injection, privilege escalation, and data leakage
  • Exceptional system-design and software-engineering fundamentals
  • Strong written and verbal communication skills, with the ability to influence technical direction across teams
  • Comfortable operating in an ambiguous, fast-moving environment with substantial ownership and autonomy
  • Passion for building production software and infrastructure rather than purely research-focused AI
Nice to Have
  • Experience with model-serving technologies such as vLLM, TensorRT-LLM, or Triton
  • Experience with Temporal, Airflow, Prefect, or similar workflow-orchestration platforms
  • Familiarity with Model Context Protocol (MCP) or other agent communication standards
  • Experience with model fine-tuning, LoRA, or quantization
  • Experience building AI infrastructure within fintech, healthcare, or another regulated industry
  • Experience working with multimodal, voice, or edge-inference systems
  • Experience designing human-in-the-loop approval and oversight systems
  • Experience building developer platforms, internal SDKs, or CI/CD automation for AI workloads
Why Apply?
  • Take Staff-level ownership over the architecture and technical direction of major AI platforms
  • Build inference and agentic infrastructure used across a global technology organization
  • Work on genuinely greenfield engineering problems spanning LLMs, autonomous agents, distributed systems, security, and cloud infrastructure
  • Remain deeply hands-on while influencing engineering standards and mentoring a high-calibre technical team
  • Join a startup-style environment with significant autonomy, backed by the scale and resources of an established global platform
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer
Staff Software Engineer

Greybridge Search & Selection • Toronto

On-site
CAD 140,000 - 175,000
Senior LLMOps Engineer -Cloud / AI Infrastructure
Senior LLMOps Engineer -Cloud / AI Infrastructure

Talent To Hire Inc. • Toronto

On-site
CAD 120,000 - 160,000
Competitive salary
Meaningful equity
Innovative work culture
Staff AI Platform Engineer - Inference & Agentic Systems
Staff AI Platform Engineer - Inference & Agentic Systems

Paytm • Toronto

On-site
CAD 140,000 - 180,000
Senior AI Engineer
Senior AI Engineer

Encore Technical Solutions Inc. • Toronto

Hybrid
CAD 140,000 - 220,000
AWS AI Solution Architect (EU/Canada)
AWS AI Solution Architect (EU/Canada)

Dedicatted» • Quebec

Hybrid
CAD 110,000 - 150,000
Staff Machine Learning Engineer – AI & Agentic Systems
Staff Machine Learning Engineer – AI & Agentic Systems

Project X Ltd. • Toronto

Hybrid
CAD 120,000 - 150,000
Staff Machine Learning Engineer – AI & Agentic Systems
Staff Machine Learning Engineer – AI & Agentic Systems

Project Limited • Toronto

On-site
CAD 120,000 - 150,000
Staff AI Platform Engineer - Inference & Agentic Systems
Staff AI Platform Engineer - Inference & Agentic Systems

Paytm Payments Services • Toronto

On-site
CAD 140,000 - 200,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank | Canada's Challenger Bank • Toronto

On-site
CAD 150,000 - 210,000
Sr. AI Engineer to develop Agentic AI solutions using Python and LangGraph for our enterprise retail client
Sr. AI Engineer to develop Agentic AI solutions using Python and LangGraph for our enterprise retail client

S I Systems • Toronto

Hybrid
CAD 120,000 - 160,000