Senior DevOps Engineer

Tru India

Toronto

On-site

CAD 100,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tru India seeks a Senior DevOps/SRE to build, operate, monitor and scale an enterprise AI platform. You will work across Dev, QA and Production, focusing on Azure infra, Kubernetes, and observability to ensure reliability and performance.

The role is hybrid with GTA residence required and occasional travel. You will design scalable CI/CD pipelines, automate infrastructure, and lead capacity planning for growth through 2027.

Qualifications

  • Experience in DevOps/SRE roles for enterprise platforms.
  • Advanced hands-on in Microsoft Azure.
  • Strong Kubernetes and Docker expertise.
  • Experience with CI/CD pipelines and IaC tools.
  • Proficient in monitoring/observability and dashboards.

Responsibilities

  • Design, deploy and maintain Azure infrastructure and platforms.
  • Manage Kubernetes clusters, containers, networking and storage.
  • Implement IaC and CI/CD automation for repeatable deployments.
  • Monitor performance, capacity and scaling of the platform.
  • Create and improve dashboards for platform visibility and reliability.
  • Provide production support and incident response planning.

Skills

DevOps
SRE
Cloud infrastructure
Kubernetes
Docker
CI/CD
Terraform
Azure
Observability
Automation

Tools

Grafana
Kibana
Azure Monitor
Application Insights
Terraform
Bicep
ARM templates
OpenTelemetry
Prometheus
Elasticsearch
Log Analytics

Job description

Position Overview

We are looking for a Senior DevOps / Site Reliability Engineer to help build, operate, monitor and scale our new enterprise AI platform. This role will be responsible for the DevOps and reliability capabilities supporting the platform across all environments including Development, QA and Production. The environment is currently focused and manageable consisting of approximately 10–30 containers but is expected to grow as the AI platform expands in 2027 and beyond.

Location

This role is hybrid, providing some flexibility to work from home. However, candidates must reside in the Greater Toronto Area (GTA) and be prepared to attend in-person meetings as required. Occasional travel to client sites or workshops may also be necessary.

Key Responsibilities
Azure Infrastructure and Platform Deployment
  • Design, deploy, configure and maintain infrastructure within Microsoft Azure.
  • Deploy and manage virtual machines, containers, Kubernetes clusters, networking, storage and supporting platform services.
  • Support infrastructure across Development, QA and Production environments.
  • Establish repeatable and reliable deployment processes using Infrastructure as Code and CI/CD automation.
  • Maintain secure, resilient and appropriately sized platform environments.
Kubernetes and Container Management
  • Deploy, configure and operate containerized applications using Kubernetes.
  • Manage container lifecycle, configuration, secrets, networking, storage and application dependencies.
  • Monitor container and cluster health, resource consumption, capacity and performance.
  • Troubleshoot deployment, networking, configuration and runtime issues.
  • Establish appropriate standards for container deployment and Kubernetes operations.
Performance, Load Management, and Scaling
  • Monitor platform demand, workload patterns, resource utilization and application performance.
  • Configure horizontal and vertical scaling policies for containers and supporting infrastructure.
  • Develop intelligent scaling approaches based on workload, queue depth, response time, resource utilization and business demand.
  • Conduct capacity planning and identify potential performance bottlenecks before they affect production.
  • Help introduce predictive or AI-assisted scaling and platform management capabilities.
Dashboards and Platform Visibility
  • Design and build advanced operational dashboards using tools such as Grafana, Kibana, Azure Monitor, Application Insights and similar technologies.
  • Create clear executive, operational, application and infrastructure views of platform health.
  • Build dashboards covering availability, performance, capacity, errors, latency, traffic, container health, AI workloads and service dependencies.
  • Establish meaningful service-level indicators, service-level objectives and reliability metrics.
  • Continuously improve dashboards so that issues, trends and risks can be quickly identified.
  • Advanced dashboard design and dashboard-building experience is a core requirement for this role.
Monitoring and Alerting
  • Implement monitoring and alerting across infrastructure, applications, containers, integrations and AI platform services.
  • Configure actionable alerts that identify real production risks while minimizing unnecessary alert noise.
  • Establish thresholds, anomaly detection, health checks, synthetic monitoring and automated remediation where appropriate.
  • Create operational runbooks and troubleshooting guidance.
  • Work with development and architecture teams to improve platform observability.
Production Reliability and Support
  • Support the stability, availability and operational readiness of the production AI platform.
  • Investigate and resolve platform, deployment, infrastructure, monitoring and performance issues.
  • Participate in root-cause analysis and implement preventative improvements.
  • Ensure that production support processes, documentation and escalation paths are established before platform usage increases.
  • Provide very light production support during 2026, with no regular after-hours support currently anticipated.
  • Help prepare the operating model for increased platform adoption and support requirements expected in 2027.
Requirements
  • Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure or platform engineering.
  • Advanced hands‑on experience with Microsoft Azure.
  • Strong experience deploying and operating Kubernetes environments.
  • Strong knowledge of containerization technologies such as Docker.
  • Experience deploying and supporting containerized applications in Development, QA and Production environments.
  • Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights or comparable tools.
  • Strong experience implementing monitoring, observability, logging, alerting and operational health checks.
  • Experience managing application load, infrastructure capacity, performance and automated scaling.
  • Experience with CI/CD pipelines and automated application deployment.
  • Experience with Infrastructure as Code tools such as Terraform, Bicep or ARM templates.
  • Strong troubleshooting skills across applications, containers, infrastructure, networking and cloud services.
  • Ability to work independently while collaborating closely with developers, architects, AI engineers and platform stakeholders.
Preferred Experience
  • Experience supporting AI, machine learning, data or high‑compute platforms.
  • Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues or model performance.
  • Experience implementing automated remediation, predictive monitoring or AI‑assisted platform operations.
  • Familiarity with AWS services and cloud operations.
  • Experience with Elasticsearch, Log Analytics, OpenTelemetry, Prometheus or similar observability technologies.
  • Experience defining service‑level indicators, service‑level objectives and reliability standards.
  • Experience with security, identity, secrets management and cloud governance within Azure.
What Makes This Role Different

This is not a large‑scale high‑pressure production support environment. The initial platform footprint is relatively focused with approximately 10–30 containers and very limited production support expected during 2026.

The role offers the opportunity to establish the platform correctly from the beginning, introduce modern DevOps and SRE practices and experiment with intelligent monitoring, automated scaling, advanced dashboards and AI‑assisted platform management. As platform adoption increases the responsibilities and operational scope are expected to grow throughout 2027.

Salary Range

$100,000-$120,000

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Operations Engineer
Senior AI Platform Operations Engineer

EQ Bank • Toronto

On-site
CAD 120,000 - 160,000
Senior DevOps Engineer
Senior DevOps Engineer

Quest Global • Vancouver

On-site
CAD 100,000 - 120,000
401(k) matching
Health insurance
Dental insurance
+5
DevOps Platform Enablement Lead
DevOps Platform Enablement Lead

PureFacts Financial Solutions Inc. • Toronto

On-site
CAD 120,000 - 160,000
Senior DevOps Engineer
Senior DevOps Engineer

Confluence Technologies, Inc. • Toronto

On-site
CAD 90,000 - 130,000
Generous time off
Career development
Social events
+1
Senior Platform Backend Developer
Senior Platform Backend Developer

Hootsuite Inc. • Vancouver

On-site
CAD 115,000 - 162,000
Senior Azure Cloud Platform Engineer
Senior Azure Cloud Platform Engineer

United States Digital Space LLC • Toronto

Hybrid
CAD 90,000 - 105,000
Senior DevOps
Senior DevOps

Quartermaster inc. • Toronto

Hybrid
CAD 160,000 - 215,000
30 days of PTO annually
Health, dental, and wellness benefits
Tech allowance benefit
+1
Senior System Administrator
Senior System Administrator

Harris Computer • Ottawa

On-site
CAD 75,000 - 85,000
Azure SRE Developer
Azure SRE Developer

Aarorn Technologies Inc • Toronto

Hybrid
Senior DevOps Engineer - Cloud-Native, Azure & AI/ML
Senior DevOps Engineer - Cloud-Native, Azure & AI/ML

Publicis Groupe Canada • Toronto

Hybrid
CAD 90,000 - 125,000