Senior Machine Learning Engineer

Jobgether

Toronto

Hybrid

CAD 185,000 - 225,000

Full time

41 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual bonus
RSU equity
Health benefits
Wellbeing days
Professional development
Hybrid work (Toronto)

Job summary

Jobgether is seeking a Senior Machine Learning Engineer based in Canada to build production infrastructure behind AI-powered products at scale. You will design and operate distributed systems that make ML and generative AI capabilities reliable, scalable, and cost-effective in production.

You will collaborate with Applied Scientists and engineers to translate research into robust customer-facing systems, shape long-term AI infrastructure direction, and tackle architecture, scalability,

Qualifications

  • 5+ years of experience building and operating production software services at scale, with strong proficiency in Python or an equivalent programming language.
  • Strong software engineering fundamentals, including system design, architecture, coding, testing, debugging, and production operations.
  • Proven experience owning production services or data pipelines, including operational or on-call responsibilities, incident response, and long-term technical debt management.
  • Deep understanding of distributed processing principles and practical experience with Spark, Dask, or comparable distributed computing technologies.
  • Strong SQL capabilities and experience working with large-scale data workloads.
  • Demonstrated experience integrating machine learning models or LLM-based capabilities into production systems, with the ability to work effectively alongside Applied Scientists or ML researchers.
  • Production experience with AWS and Kubernetes, including deploying and operating cloud-native workloads.
  • Familiarity with machine learning technologies such as MLFlow, TensorFlow, or PyTorch and data orchestration tools such as Airflow or Prefect is advantageous.

Responsibilities

  • Lead the design and implementation of large-scale, production-grade distributed systems that support AI and machine learning features used by millions of users.
  • Shape the longer-term technical vision for AI infrastructure in collaboration with staff and senior staff engineers, translating strategic direction into practical, deliverable initiatives.
  • Make architecture decisions that balance scalability, reliability, flexibility, operational simplicity, and cost effectiveness.
  • Own production services and pipelines, including operational health, on-call responsibilities, incident response, monitoring, and technical debt management.
  • Build infrastructure and engineering interfaces that enable Applied Scientists to safely and reliably transition machine learning and LLM models from research into production.
  • Develop and operate scalable data workloads using Python, SQL, and distributed processing technologies such as Spark or Dask.
  • Deploy and maintain production systems across AWS and Kubernetes environments, ensuring they meet appropriate reliability and performance standards.
  • Integrate production-ready generative AI and large language model capabilities into customer-facing product experiences.
  • Improve data usability and engineering practices across the AI Products organization, reducing operational toil and raising overall technical quality.
  • Provide technical leadership and mentorship to engineers, helping raise engineering standards and supporting the development of less-experienced team members.
  • Collaborate across engineering, data, science, and product teams to drive technical initiatives, resolve complex problems, and build consensus around architectural decisions.
  • Identify and address technical debt, infrastructure risks, and opportunities to improve the scalability and maintainability of the AI technology estate.

Skills

Python
Distributed systems
SQL
AWS
Kubernetes
Spark
Dask
LLMs
ML frameworks

Tools

Spark
Dask
Airflow
Prefect
TensorFlow
PyTorch
MLFlow
AWS
Kubernetes

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for aSenior Machine Learning Engineerbased inCanada.

This is a senior engineering opportunity focused on building the production infrastructure behind AI-powered products used at significant scale.
You will design and operate distributed systems that make machine learning and generative AI capabilities reliable, scalable, and cost-effective in production.
Working alongside Applied Scientists and engineers, you will turn advanced models and research into robust customer-facing systems.
You will help shape the long-term technical direction of AI infrastructure while tackling complex architecture, scalability, reliability, and technical debt challenges.
The role combines hands-on software engineering with technical leadership, mentoring, and cross-functional collaboration.
You will work with modern technologies including Python, SQL, Spark or Dask, AWS, Kubernetes, and emerging LLM technologies.
This is an ideal role for an experienced engineer who enjoys solving complex distributed-systems problems and making AI capabilities dependable at scale.

Accountabilities
  • Lead the design and implementation of large-scale, production-grade distributed systems that support AI and machine learning features used by millions of users.
  • Shape the longer-term technical vision for AI infrastructure in collaboration with staff and senior staff engineers, translating strategic direction into practical, deliverable initiatives.
  • Make architecture decisions that balance scalability, reliability, flexibility, operational simplicity, and cost effectiveness.
  • Own production services and pipelines, including operational health, on-call responsibilities, incident response, monitoring, and technical debt management.
  • Build infrastructure and engineering interfaces that enable Applied Scientists to safely and reliably transition machine learning and LLM models from research into production.
  • Develop and operate scalable data workloads using Python, SQL, and distributed processing technologies such as Spark or Dask.
  • Deploy and maintain production systems across AWS and Kubernetes environments, ensuring they meet appropriate reliability and performance standards.
  • Integrate production-ready generative AI and large language model capabilities into customer-facing product experiences.
  • Improve data usability and engineering practices across the AI Products organization, reducing operational toil and raising overall technical quality.
  • Provide technical leadership and mentorship to engineers, helping raise engineering standards and supporting the development of less-experienced team members.
  • Collaborate across engineering, data, science, and product teams to drive technical initiatives, resolve complex problems, and build consensus around architectural decisions.
  • Identify and address technical debt, infrastructure risks, and opportunities to improve the scalability and maintainability of the AI technology estate.
Requirements
  • 5+ years of experience building and operating production software services at scale, with strong proficiency in Python or an equivalent programming language.
  • Strong software engineering fundamentals, including system design, architecture, coding, testing, debugging, and production operations.
  • Proven experience owning production services or data pipelines, including operational or on-call responsibilities, incident response, and long-term technical debt management.
  • Deep understanding of distributed processing principles and practical experience with Spark, Dask, or comparable distributed computing technologies.
  • Strong SQL capabilities and experience working with large-scale data workloads.
  • Demonstrated experience integrating machine learning models or LLM-based capabilities into production systems, with the ability to work effectively alongside Applied Scientists or ML researchers.
  • Production experience with AWS and Kubernetes, including deploying and operating cloud-native workloads.
  • Familiarity with machine learning technologies such as MLFlow, TensorFlow, or PyTorch and data orchestration tools such as Airflow or Prefect is advantageous.
  • Prior experience applying or fine-tuning LLMs in a product environment is a plus, but strong production engineering expertise remains the primary requirement.
  • Strong technical leadership skills, with the ability to set direction, make sound architectural decisions, mentor engineers, and raise engineering standards.
  • Excellent communication and collaboration skills, particularly when working across multidisciplinary teams and translating complex technical concepts into practical decisions.
  • A proactive, pragmatic approach to problem solving, with the ability to navigate ambiguity and drive meaningful technical outcomes.
  • Willingness to participate in an on‑call rotation and take ownership of the reliability of production systems.
  • A growth‑oriented mindset, curiosity about emerging AI technologies, and enthusiasm for applying new approaches responsibly in production environments.
Benefits
  • Annual base salary range ofCA$185,000–CA$225,000, with individual compensation determined by factors such as geography, experience, skills, and role level.
  • Eligibility for annual performance bonuses and equity through RSU programs for permanent employees.
  • Potential access to additional performance-based cash or equity incentives depending on role level and company performance.
  • Comprehensive health, wellness, and retirement programs.
  • Wellbeing days and generous paid leave.
  • Dedicated professional development budgets to support ongoing learning and career growth.
  • Flexible hybrid working model combining remote autonomy with access to modern office spaces in Toronto.
  • Collaborative “boost days” designed to support team connection, knowledge sharing, and effective delivery.
  • Opportunity to work on AI-powered products and production infrastructure serving millions of users.
  • A multidisciplinary environment bringing together engineers, Applied Scientists, product managers, analysts, and data specialists.
  • A culture that values skills, impact, curiosity, continuous learning, and diverse perspectives.
  • Inclusive and accessible workplace practices, with support and reasonable accommodations available throughout the hiring process.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff+ Software Engineer, AI
Staff+ Software Engineer, AI

Jobgether • Canada

On-site
CAD 170,000 - 260,000
Fully remote across the Americas
Equity participation
Flexible hours
+4
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Motion Recruitment Partners LLC • Toronto

On-site
CAD 180,000 - 280,000
Lead Product Manager - AI/ML
Lead Product Manager - AI/ML

STACK IT Recruitment • Toronto

Hybrid
CAD 130,000 - 155,000
Health & Wellness Benefits
RRSP Matching
Annual Bonus
+1
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Motion Recruitment • Toronto

On-site
CAD 140,000 - 190,000
Bonus eligible
Medical, Dental, Vision Insurance
Vacation Time
AI/ML Infrastructure Engineer
AI/ML Infrastructure Engineer

Avenue Code • Toronto

Hybrid
CAD 165,000 - 175,000
Machine Learning Engineer
Machine Learning Engineer

Invision AI • Toronto

Hybrid
CAD 110,000 - 150,000
Equity
RRSP Plan
Health and Dental
+1
Principal Machine Learning Engineer
Principal Machine Learning Engineer

Equinix • Toronto

On-site
CAD 154,000 - 232,000
Employee Assistance Program
Canada Core Benefits
Machine Learning Engineer , Amazon Customer Service
Machine Learning Engineer , Amazon Customer Service

Amazon • Vancouver

On-site
CAD 115,000 - 192,000
Health insurance
RRSP
DPSP
+1
Sr. AI Engineer
Sr. AI Engineer

TheAppLabb • Toronto

On-site
CAD 110,000 - 170,000
Competitive salary
Opportunities for career growth
Fitness challenge incentives
AI Software Engineer
AI Software Engineer

Talentlab • Canada

Remote
CAD 90,000 - 130,000