Technical Architect

Mphasis

Toronto

On-site

CAD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mphasis in Canada is seeking an experienced MLOps Engineer to design, deploy and monitor end-to-end ML/LLM pipelines using open-source tools and cloud-native services across AWS, GCP, and Azure.

You will implement CI/CD for ML models, manage data/versioning with Mlflow, Kubeflow, DVC, and Airflow, and ensure performance, observability, and cost-efficiency of models in production. The role requires hands-on expertise and collaboration across teams.

Qualifications

  • Hands-on experience deploying ML/LLM pipelines and applications.
  • Experience with open‑source MLOps tools and cloud-native ML services.
  • Strong cloud platform knowledge across AWS, GCP or Azure.

Responsibilities

  • Design, build, and manage MLOPs and LLMOPs pipelines for data ingestion, model training, validation, deployment, and monitoring.
  • Set up cloud-native MLOPs pipelines on AWS, GCP and Azure and manage versioning, retraining, and deployment workflows.
  • Implement CI/CD pipelines for ML models using GitHub Actions, Jenkins, or GitLab CI/CD.
  • Monitor production models with observability tools and optimize performance and cost.

Skills

CI/CD practices
Python
Bash scripting
DevOps methodologies

Education

Bachelor's/Master's in CS/Engineering or related field

Tools

Kubernetes
Docker
Mlflow
Kubeflow
DVC
Airflow
GitHub Actions
Jenkins
GitLab CI
Terraform
Ansible
AWS
GCP
Azure

Job description

MLOps Engineer (Cloud Engineer + DevOps)

Location: Canada

Job Overview

We are looking for an experienced MLOPs / LLMOPs Engineer with a strong background in deploying and monitoring machine learning and large language model (LLM) pipelines. The ideal candidate will have 10+ years of experience in MLOPs, with expertise in setting up end‑to‑end ML/LLM pipelines using open‑source tools and cloud‑native solutions on platforms like AWS, GCP, and Azure. This role requires hands‑on knowledge in deploying, automating, and monitoring ML/LLM workflows, with a solid grounding in DevOps practices to ensure seamless CI/CD processes.

Responsibilities
  • Pipeline Design & Implementation:
    • Design, build, and manage MLOPs and LLMOPs pipelines for data ingestion, model training, validation, deployment, and monitoring.
    • Use open‑source tools such as Mlflow, Kubeflow, DVC, and Airflow to automate and monitor machine learning workflows.
    • Implement scalable LLM‑specific solutions for model training and inference, optimizing resource allocation and deployment efficiency.
  • Cloud‑native MLOPs Implementation:
    • Set up and manage MLOPs pipelines in primary GCP (SageMaker, EKS, Lambda, S3), or have similar experience with GCP (Vertex AI, AI Platform Pipelines), and Azure (Machine Learning, AKS, Azure Functions).
    • Manage model versioning, retraining, and deployment workflows on cloud platforms to ensure consistent performance and availability.
    • Execute CI/CD pipelines for ML models with GitHub Actions, Jenkins, or GitLab CI.
  • Model Monitoring & Performance Optimization:
    • Monitor models in production using Prometheus, Grafana, and Tensorboard, establishing observability metrics for model drift, accuracy, and latency.
    • Collaborate with Data Engineering and ML teams to implement scalable and efficient pipelines.
    • Use A/B testing and shadow deployment strategies to validate and optimize LLM model performance in real‑time.
  • LLM‑specific Model Operations:
    • Deploy and monitor LLMs for specific tasks, ensuring they adhere to performance SLAs and are optimized for cost.
    • Understand techniques of fine‑tuning, optimizing inference, and managing infrastructure costs for large LLMs.
Required Skills and Qualifications
  • Technical Skills – Good to have:
    • Proficiency with Kubernetes and Docker for container orchestration and model deployment.
    • Experience with open‑source MLOPs tools (Mlflow, Kubeflow, DVC) and data versioning.
    • Hands‑on experience with cloud‑native ML tools in AWS, GCP, or Azure and associated ML services.
    • Knowledge of Python or Bash scripting for automating processes and custom integrations.
  • DevOps-Related Skills:
    • Solid understanding of CI/CD practices and tools like GitHub Actions, Jenkins, or GitLab CI/CD to build and deploy ML/LLM models.
    • Proficient in infrastructure‑as‑code tools, such as Terraform or Ansible, to enable automated provisioning and configuration management.
  • Programming & Scripting:
    • Python
    • SQL, No‑SQL, PySpark (Optional)
  • AI/ML & Data Science – Good To Have:
    • Supervised, Unsupervised Learning & Model evaluation metrics
    • NLP, RAG, GenAI, LLMs
    • Deep Learning (Sequential & Functional APIs) using Pytorch/TensorFlow
    • MLOPs & Mlflow Experiment Tracking
    • Explainable AI (XAI) LIME SHAP (Optional)
  • Cloud Platforms:
    • Primary: AWS or hands‑on expertise on any of cloud platforms
    • Azure AI/ML
    • Google Vertex AI
    • Databricks Studio
  • Education:
    • Bachelors/ Master's/ PhD Degree in Mathematics, Statistics, Physics, Computer Science, Engineering, Data Science, or a related quantitative field.
  • Process Skills:
    • Understanding of Agile and Scrum methodologies.
    • Ability to follow SDLC processes and contribute to technical documentation.
  • Behavioral Skills:
    • Structural thinking and goal‑oriented approach to problem‑solving.
    • Self‑motivated and capable of working independently with minimal management supervision.
    • Well‑developed design, analytical & problem‑solving skills.
    • Excellent communication and interpersonal skills.
    • Excellent team player, able to work with virtual teams in several time zones.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Ops Developer
Machine Learning Ops Developer

Autodesk • Toronto

On-site
CAD 80,000 - 120,000
ML Platform Engineer – Google Cloud (GCP) and Vertex AI
ML Platform Engineer – Google Cloud (GCP) and Vertex AI

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Mississauga

On-site
CAD 80,000 - 110,000
Data Engineer
Data Engineer

CoFoMo Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 140,000
Junior Machine Learning Engineer
Junior Machine Learning Engineer

Jobtailor • North Vancouver

On-site
CAD 90,000 - 140,000
Senior LLMOps Engineer -Cloud / AI Infrastructure
Senior LLMOps Engineer -Cloud / AI Infrastructure

TEEMA Solutions Group • Toronto

Hybrid
CAD 120,000 - 160,000
Competitive salary
Meaningful equity
Innovative work culture
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Tundra Technical Solutions • Toronto

On-site
CAD 140,000 - 190,000
Senior LLMOps
Senior LLMOps

Marler & Associates Search • Quebec

Remote
CAD 120,000 - 180,000
Stock options
Flexible remote work
Focus on long-term career growth
Gen AI Engineer
Gen AI Engineer

Smart IT Frame LLC • Halifax

On-site
CAD 100,000 - 150,000
Senior Software Engineer, AI Products
Senior Software Engineer, AI Products

HRB • Kitchener

On-site
CAD 110,000 - 150,000
Data/AI Engineer
Data/AI Engineer

Jobtailor • Vancouver

On-site
CAD 110,000 - 150,000