Senior LLMOps Engineer -Cloud / AI Infrastructure

TEEMA Solutions Group

Toronto

On-site

CAD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Meaningful equity
Innovative work culture

Job summary

A rapid-growth technology firm in Toronto is seeking a Staff LLMOps Engineer to lead the design and optimization of large language model infrastructure on the cloud. The ideal candidate has over 6 years of experience in DevOps and expertise in deploying LLMs in cloud environments. Responsibilities include architecting deployment pipelines and ensuring high-performance AI applications. Competitive salary and equity are included in the offer.

Qualifications

  • 6+ years in DevOps or cloud platform engineering.
  • 2+ years of experience with LLMs deployment.
  • Expertise with GPU-accelerated inference.

Responsibilities

  • Architect and operationalize LLM deployment pipelines on AWS.
  • Build and scale multi-GPU inference infrastructure.
  • Optimize inference performance using various frameworks.

Skills

DevOps expertise
ML infrastructure knowledge
Cloud platform experience
Python proficiency
Monitoring tools integration

Tools

AWS
Kubernetes
Terraform
Prometheus
Grafana

Job description

Location: Downtown Toronto
Hybrid: 4 days in office

Ready to build what powers the next generation of AI?

We’re looking for a Staff LLMOps Engineer to lead the design, deployment, and optimization of large language model (LLM) infrastructure on the cloud.
You’ll be the driving force behind taking trained models from lab to production—scaling efficiently across multi-GPU clusters and pushing the boundaries of inference performance for enterprise-grade AI applications.

If you thrive at the intersection of AI, cloud engineering, and systems optimization, this is your chance to shape the future of large-scale model serving in a high-impact environment.

What You’ll Do

Architect and operationalize LLM deployment pipelines on AWS and Kubernetes/EKS.

Build and scale multi-GPU inference infrastructure for low latency, high availability, and cost efficiency.

Optimize inference using frameworks like vLLM, SGLang, and DeepSpeed-Inference.

Implement advanced serving techniques: continuous batching, speculative decoding, KV-cache management, and distributed scheduling.

Collaborate with AI researchers to convert model training outputs into production-grade APIs and services.

Establish observability and monitoring for latency, throughput, GPU utilization, and failure recovery.

Automate provisioning, scaling, and upgrades using Terraform and CI/CD pipelines.

Ensure compliance, security, and efficiency in multi-tenant LLM hosting for enterprise clients.

What We’re Looking For

6+ years in DevOps, ML infrastructure, or cloud platform engineering.

2+ years of direct experience deploying and optimizing LLMs or large-scale ML models.

Expertise with GPU-accelerated inference and distributed serving environments.

Deep familiarity with cloud-native architectures (AWS, GCP, Azure) and Kubernetes.

Strong foundation in Python, Bash, and IaC (Terraform).

Experience integrating monitoring tools (Prometheus, Grafana, Datadog) for performance visibility.

Passion for building robust, scalable, and secure AI systems.

Why Join

Lead and own mission-critical AI infrastructure at a fast-scaling startup.

Work alongside world-class engineers, data scientists, and innovators.

Competitive salary + meaningful equity in a company redefining applied AI.

A culture built on innovation, technical depth, and impact—your work truly matters.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Encore Technical Solutions Inc. • Toronto

On-site
CAD 100,000 - 150,000
Senior LLMOps Engineer — Cloud AI Inference + Equity
Senior LLMOps Engineer — Cloud AI Inference + Equity

TEEMA Solutions Group • Toronto

Hybrid
CAD 120,000 - 160,000
Competitive salary
Meaningful equity
Innovative work culture
Senior LLMOps
Senior LLMOps

Marler & Associates Search • Quebec

Remote
CAD 120,000 - 180,000
Stock options
Flexible remote work
Focus on long-term career growth
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Motion Recruitment • Toronto

On-site
CAD 140,000 - 190,000
Bonus eligible
Medical, Dental, Vision Insurance
Vacation Time
AI/ML Platform Engineer (Azure & GCP)
AI/ML Platform Engineer (Azure & GCP)

IT Accel, Inc. • Toronto

Hybrid
CAD 140,000 - 180,000
AVP, Lead AI Engineer | Principal Engineer
AVP, Lead AI Engineer | Principal Engineer

Encore Technical Solutions Inc. • Toronto

On-site
CAD 180,000 - 260,000
DL Performance Software Engineer - LLM Inference
DL Performance Software Engineer - LLM Inference

United States Digital Space LLC • Toronto

Hybrid
CAD 135,000 - 220,000
Equity
Benefits
Sr. AI Engineer
Sr. AI Engineer

TheAppLabb • Toronto

On-site
CAD 110,000 - 170,000
Competitive salary
Opportunities for career growth
Fitness challenge incentives
Technical Architect
Technical Architect

Mphasis • Toronto

On-site
CAD 120,000 - 160,000
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency

Cohere • Montreal

Remote
CAD 100,000 - 130,000
Open and inclusive culture
Cutting-edge AI research collaboration
Weekly lunch stipend
+5