As an AI ModelOps Engineer, you will be part of the Enterprise AI Platforms & AI ModelOps team within the AI and Data group at Canadian Tire Corporation (CTC). In this role, you will support the infrastructure and provisioning of environments that enable AI teams to build and deploy agentic AI solutions. You will help design, implement, operate, and optimize the platforms, tools, and processes that support AI solutions across their lifecycle.
You will contribute to the evolution of CTC's AAAI platform and MLOps ecosystem, enabling the development, deployment, monitoring, and governance of machine learning solutions at scale. You will also help establish and mature CTC's Agentic AI platform capabilities, including support for generative AI, large language models (LLMs), AI agents, agentic workflows, and emerging AI engineering practices.
Responsibilities
- Design, implement, operate, and continuously improve platforms, tools, and processes supporting the end-to-end lifecycle of AI solutions
- Enable deployment, monitoring, governance, and lifecycle management of ML models, generative AI solutions, and AI agents in production
- Contribute to the evolution and operational excellence of CTC's AAAI platform and MLOps ecosystem
- Help establish and mature Agentic AI platform capabilities including infrastructure, tooling, registries, observability, and evaluation frameworks
- Develop and maintain automation, CI/CD pipelines, and platform services that accelerate delivery of AI solutions
- Design and implement observability, monitoring, alerting, and troubleshooting capabilities for AI platforms
- Collaborate with data scientists, AI engineers, cloud engineers, architects, and IT teams
- Evaluate, recommend, and implement emerging AI platform technologies and engineering practices
- Support platform governance including model and agent lifecycle management, auditability, and responsible AI practices
- Contribute to platform standards, reusable components, reference architectures, and best practices
- Ensure availability, scalability, resilience, security, and cost efficiency of AI platform infrastructure
- Perform post-deployment analysis and operational reviews
- Stay current with industry trends in ML, generative AI, Agentic AI, and cloud platforms
Requirements
- Hands-on experience with MLOps, GenAI operations, or AI platform engineering
- Experience building and operating AI, ML, or data platforms in cloud environments, preferably Microsoft Azure
- Strong understanding of machine learning concepts, model lifecycle management, and governance
- Experience with generative AI, LLMs, RAG, AI agents, and emerging AI engineering practices
- Proficiency in Python and experience developing AI-enabled applications
- Experience with Azure AI Foundry, Azure Machine Learning, Databricks, MLflow, or similar platforms
- Experience with containerization: Docker, Kubernetes, and related orchestration platforms
- Familiarity with DevOps practices including CI/CD pipelines, source control, automation, and IaC
- Practical experience with Terraform, Bicep, or equivalent IaC tools
- Knowledge of cloud security, governance, access management, and responsible AI practices
- Excellent verbal and written communication skills
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related discipline
Bachelor's degree in Computer Science, Engineering, Information Technology, or related discipline, or equivalent combination of education, certifications, and experience. Hands-on MLOps or AI platform engineering experience. Proficiency in Python. Experience with Azure, Docker, Kubernetes, Terraform or Bicep, and AI platforms such as Azure AI Foundry or Databricks.