Senior ML Infrastructure Engineer for AI at Scale

Jobgether

Toronto

Hybrid

CAD 185,000 - 225,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual bonus
RSU equity
Health benefits
Wellbeing days
Professional development
Hybrid work (Toronto)

Job summary

Jobgether is seeking a Senior Machine Learning Engineer based in Canada to build production infrastructure behind AI-powered products at scale. You will design and operate distributed systems that make ML and generative AI capabilities reliable, scalable, and cost-effective in production.

You will collaborate with Applied Scientists and engineers to translate research into robust customer-facing systems, shape long-term AI infrastructure direction, and tackle architecture, scalability,

Qualifications

  • 5+ years of experience building and operating production software services at scale, with strong proficiency in Python or an equivalent programming language.
  • Strong software engineering fundamentals, including system design, architecture, coding, testing, debugging, and production operations.
  • Proven experience owning production services or data pipelines, including operational or on-call responsibilities, incident response, and long-term technical debt management.
  • Deep understanding of distributed processing principles and practical experience with Spark, Dask, or comparable distributed computing technologies.
  • Strong SQL capabilities and experience working with large-scale data workloads.
  • Demonstrated experience integrating machine learning models or LLM-based capabilities into production systems, with the ability to work effectively alongside Applied Scientists or ML researchers.
  • Production experience with AWS and Kubernetes, including deploying and operating cloud-native workloads.
  • Familiarity with machine learning technologies such as MLFlow, TensorFlow, or PyTorch and data orchestration tools such as Airflow or Prefect is advantageous.

Responsibilities

  • Lead the design and implementation of large-scale, production-grade distributed systems that support AI and machine learning features used by millions of users.
  • Shape the longer-term technical vision for AI infrastructure in collaboration with staff and senior staff engineers, translating strategic direction into practical, deliverable initiatives.
  • Make architecture decisions that balance scalability, reliability, flexibility, operational simplicity, and cost effectiveness.
  • Own production services and pipelines, including operational health, on-call responsibilities, incident response, monitoring, and technical debt management.
  • Build infrastructure and engineering interfaces that enable Applied Scientists to safely and reliably transition machine learning and LLM models from research into production.
  • Develop and operate scalable data workloads using Python, SQL, and distributed processing technologies such as Spark or Dask.
  • Deploy and maintain production systems across AWS and Kubernetes environments, ensuring they meet appropriate reliability and performance standards.
  • Integrate production-ready generative AI and large language model capabilities into customer-facing product experiences.
  • Improve data usability and engineering practices across the AI Products organization, reducing operational toil and raising overall technical quality.
  • Provide technical leadership and mentorship to engineers, helping raise engineering standards and supporting the development of less-experienced team members.
  • Collaborate across engineering, data, science, and product teams to drive technical initiatives, resolve complex problems, and build consensus around architectural decisions.
  • Identify and address technical debt, infrastructure risks, and opportunities to improve the scalability and maintainability of the AI technology estate.

Skills

Python
Distributed systems
SQL
AWS
Kubernetes
Spark
Dask
LLMs
ML frameworks

Tools

Spark
Dask
Airflow
Prefect
TensorFlow
PyTorch
MLFlow
AWS
Kubernetes

Job description

Jobgether is seeking a Senior Machine Learning Engineer based in Canada to build production infrastructure behind AI-powered products at scale. You will design and operate distributed systems that make ML and generative AI capabilities reliable, scalable, and cost-effective in production.

You will collaborate with Applied Scientists and engineers to translate research into robust customer-facing systems, shape long-term AI infrastructure direction, and tackle architecture, scalability,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer: Scalable AI Infra & Orchestration
Senior ML Engineer: Scalable AI Infra & Orchestration

HelloFresh • Toronto

Hybrid
CAD 170,000 - 190,000
Box discounts
Health & dental benefits
Generous vacation & PTO
+5
Senior Production AI Systems Engineer
Senior Production AI Systems Engineer

Jaide Health • Toronto

Hybrid
CAD 140,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+5
Senior ML Engineer - Remote Canada, AI Platform & MLOps
Senior ML Engineer - Remote Canada, AI Platform & MLOps

Hyatt • Canada

On-site
CAD 90,000 - 110,000
Annual hotel stays
Flexible work schedule
RRSP matching
+2
Staff GenAI ML Engineer — Secure, Scalable AI Systems
Staff GenAI ML Engineer — Secure, Scalable AI Systems

Triwill Group • Toronto

Hybrid
CAD 168,000 - 231,000
Senior ML Engineer - Build Production-Grade AI Systems
Senior ML Engineer - Build Production-Grade AI Systems

TheAppLabb • Toronto

On-site
CAD 110,000 - 170,000
Competitive salary
Opportunities for career growth
Fitness challenge incentives
Senior ML Engineer: Agentic AI & Cloud Orchestration
Senior ML Engineer: Agentic AI & Cloud Orchestration

TEEMA • Canada

On-site
CAD 90,000 - 120,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Motion Recruitment Partners LLC • Toronto

On-site
CAD 180,000 - 280,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Motion Recruitment • Toronto

On-site
CAD 140,000 - 190,000
Bonus eligible
Medical, Dental, Vision Insurance
Vacation Time
MLOps Field Engineer: Hands-on AI Infra, Global Travel
MLOps Field Engineer: Hands-on AI Infra, Global Travel

Jobgether • Canada

On-site
CAD 90,000 - 130,000
Learning budget USD 2,000 per year
Travel opportunities
Annual compensation review
Senior AI/ML Engineer: Production GenAI & MLOps
Senior AI/ML Engineer: Production GenAI & MLOps

Socket.dev • Ottawa

Hybrid
CAD 85,000 - 129,000
Comprehensive Benefits
Hybrid work options
Professional development