Technical Architect

Mphasis

Toronto

On-site

CAD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mphasis is seeking an experienced MLOps Engineer to design, deploy, and monitor ML/LLM pipelines using open-source tools across AWS, GCP, and Azure. You will work on end-to-end ML workflows, ensuring robust CI/CD practices and scalable cloud deployments.

The role emphasizes hands-on implementation of Kubernetes, Docker, and ML tooling, with a strong DevOps focus and cross-team collaboration across time zones.

Qualifications

  • Extensive experience (10+ years) building and deploying ML/LLM pipelines.
  • Hands-on with open-source MLOps tools and cloud ML services.
  • Strong DevOps and CI/CD practices for ML workflows.

Responsibilities

  • Design, build, and monitor MLOPs and LLMOPs pipelines for data ingestion, model training, validation, deployment, and monitoring.
  • Implement open-source tooling to automate ML workflows.
  • Set up cloud-native ML pipelines on AWS, GCP, and Azure and ensure CI/CD compatibility.

Skills

Python
CI/CD mindset
Bash scripting

Education

Bachelor's/Master's/PhD in quantitative field

Tools

Kubernetes
Docker
Mlflow
Kubeflow
DVC
Airflow
SageMaker
Vertex AI
Azure ML

Job description

MLOps Engineer (Cloud Engineer + DevOps)
Location: Canada
Job Overview

We are looking for an experienced

POSITION / TITLE:
MLOps Engineer (Cloud Engineer + DevOps)
Location: Canada
Job Overview

We are looking for an experienced MLOPs / LLMOPs Engineer with a strong background in deploying and monitoring machine learning and large language model (LLM) pipelines. The ideal candidate will have 10+ years of experience in MLOPs, with expertise in setting up end-to-end ML/LLM pipelines using open-source tools and cloud-native solutions on platforms like AWS, GCP, and Azure. This role requires hands-on knowledge in deploying, automating, and monitoring ML/LLM workflows, with a solid grounding in DevOps practices to ensure seamless CI/CD processes.

Responsibilities
  • Pipeline Design & Implementation:
    • Design, build, and manage MLOPs and LLMOPs pipelines for data ingestion, model training, validation, deployment, and monitoring.
    • Use open-source tools such as Mlflow, Kubeflow, DVC, and Airflow to automate and monitor machine learning workflows.
    • Implement scalable LLM-specific solutions for model training and inference, optimizing resource allocation and deployment efficiency.
  • Cloud-native MLOPs Implementation:
    • Set up and manage MLOPs pipelines in Primary in GCP (SageMaker, EKS, Lambda, S3), or have similar experience with GCP (Vertex AI, AI Platform Pipelines), and Azure (Machine Learning, AKS, Azure Functions).
    • Manage model versioning, retraining, and deployment workflows on cloud platforms to ensure consistent performance and availability.
    • Execute CI/CD pipelines for ML models with GitHub Actions, Jenkins, or GitLab CI.
  • Model Monitoring & Performance Optimization:
    • Monitor models in production using Prometheus, Grafana, and Tensorboard, establishing observability metrics for model drift, accuracy, and latency.
    • Collaborate with Data Engineering and ML teams to implement scalable and efficient pipelines
    • Use A/B testing and shadow deployment strategies to validate and optimize LLM model performance in real-time.
  • LLM-specific Model Operations:
    • Deploy and monitor LLMs for specific tasks, ensuring they adhere to performance SLAs and are optimized for cost.
    • Understand techniques of fine-tuning, optimizing inference, and managing infrastructure costs for large LLMs.
Required Skills And Qualifications
Technical Skills – Good to have:
  • Proficiency with Kubernetes and Docker for container orchestration and model deployment.
  • Experience with open-source MLOPs tools (Mlflow, Kubeflow, DVC) and data versioning.
  • Hands-on experience with cloud-native ML tools in AWS, GCP, or Azure and associated ML services.
  • Knowledge of Python or Bash scripting for automating processes and custom integrations.
DevOps-Related Skills
  • Solid understanding of CI/CD practices and tools like GitHub Actions, Jenkins, or GitLab CI/CD to build and deploy ML/LLM models.
  • Proficient in infrastructure-as-code tools, such as Terraform or Ansible, to enable automated provisioning and configuration management.
Programming & Scripting
  • Python
  • SQL, No-SQL, PySpark (Optional)
AI/ML & Data Science - Good To Have
  • Supervised, Unsupervised Learning & Model evaluation metrics
  • NLP, RAG, GenAI, LLMs
  • Deep Learning (Sequential & Functional APIs) using Pytorch/TensorFlow
  • MLOPs & Mlflow Experiment Tracking
  • Explainable AI (XAI) LIME SHAP(Optional)
Cloud Platforms
  • (Primary: AWS) or Handons Expertise on any of cloud platforms
  • Azure AI/ML,
  • Google Vertex AI,
  • Databricks Studio
Education
  • Bachelors/ Master’s/PhD Degree in Mathematics, Statistics, Physics, Computer Science, Engineering, Data Science, or a related relevant degree from quantitative field.
Process Skills
  • Understanding ofAgile and Scrummethodologies.
  • Ability to follow SDLC processes and contribute to technical documentation.
Behavioral Skills
  • Must have Structural thinking and goal-oriented approach to problem-solving
  • Self-motivated and capable of working independently with minimal management supervision.
  • Well-developed design, analytical & problem-solving skills
  • Excellent communication and interpersonal skills.
  • Excellent team player, able to work with virtual teams in several time zones.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLMOps Engineer -Cloud / AI Infrastructure
Senior LLMOps Engineer -Cloud / AI Infrastructure

Talent To Hire Inc. • Toronto

On-site
CAD 120,000 - 160,000
Competitive salary
Meaningful equity
Innovative work culture
ML Platform Engineer – Google Cloud (GCP) and Vertex AI
ML Platform Engineer – Google Cloud (GCP) and Vertex AI

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Mississauga

On-site
CAD 80,000 - 110,000
Data Engineer
Data Engineer

CoFoMo Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 140,000
Full-Stack AI Developer
Full-Stack AI Developer

CoFoMo Inc. • Montreal (administrative region)

Hybrid
CAD 95,000 - 140,000
Senior MLOps Engineer
Senior MLOps Engineer

Deep Genomics • Toronto

On-site
CAD 175,000 - 200,000
Stock options
Comprehensive benefits (health/vision)
Flexible work environment
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Tundra Technical Solutions • Toronto

On-site
CAD 140,000 - 190,000
AI/ML Engineer
AI/ML Engineer

BrainWave Professionals • Canada

On-site
CAD 120,000 - 160,000
Python Developer
Python Developer

AIT Global inc. • Mississauga

On-site
CAD 110,000 - 180,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Hallmark Global Solutions Ltd • Mississauga

On-site
CAD 100,000 - 150,000
Python developer
Python developer

I8IS - Infiniti Software Solutions • Mississauga

On-site
CAD 120,000 - 160,000