Site Reliability Engineer

TJX Canada – Winners, Marshalls, HomeSense

Mississauga

On-site

CAD 88,000 - 123,000

Full time

17 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Associate discounts
Health care spending account
Retirement savings program
Assistance program
Vacation + tuition reimbursement
Training and development resources

Job summary

TJX Canada in Mississauga is seeking a Site Reliability Engineer to drive intelligent operations by leveraging AIOps, GenAI, automation, and self‑healing technologies across enterprise platforms. You’ll shape reliability practices and accelerate incident response while reducing manual effort through automation.

You’ll work with cutting‑edge AI technologies to enhance scalability and performance of critical business applications, collaborating across teams to foster operational excellence in a

Qualifications

  • 6+ years of experience in SRE, DevOps, Production Support Engineering, Cloud Operations, Automation Engineering, or related roles.
  • Hands‑on experience with observability/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
  • Strong coding/scripting experience in Python, Java, PowerShell, Go, or Shell scripting.
  • Hands‑on experience with automation and DevOps tools such as Azure DevOps, GitHub Actions, Jenkins, Ansible, Terraform, Power Automate, Rundeck, or similar tools.

Responsibilities

  • Implement SRE best practices across production support, incident management, problem management, change management, release, and deployment processes.
  • Build and maintain automation scripts, runbooks, self‑healing workflows, and operational tools to reduce manual effort and improve MTTR.
  • Support AI‑driven operational capabilities such as incident summarization, alert enrichment, anomaly detection, log analysis, event correlation, and root cause recommendations.
  • Configure and enhance observability across logs, metrics, traces, dashboards, alerts, and synthetic monitoring.
  • Work with monitoring/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
  • Support implementation of SLIs, SLOs, SLAs, service health metrics, reliability KPIs, and operational dashboards.
  • Integrate monitoring, ITSM, CI/CD, cloud, and automation platforms using APIs, scripts, and workflows.
  • Participate in incident response, troubleshooting, RCA, postmortems, and corrective/preventive action planning.
  • Drive shift‑left reliability by embedding monitoring, alerting, automation, and operational readiness into SDLC and CI/CD pipelines.
  • Support reliable application and platform operations in Azure Cloud.
  • Create and maintain documentation, runbooks, knowledge articles, and automation playbooks.
  • Ensure AI‑enabled operations follow enterprise security, compliance, data privacy, and responsible‑AI guidelines.

Skills

Python
Java
PowerShell
Go
Shell scripting

Tools

Azure DevOps
GitHub Actions
Jenkins
Ansible
Terraform
Power Automate
Rundeck

Job description

TJX Companies

At TJX Canada, every day brings new opportunities for growth, exploration, and achievement. You’ll be part of our vibrant team that embraces diversity, fosters collaboration, and prioritizes your development. Whether you’re working in our Distribution Centers, Corporate Offices, or Retail Stores—WINNERS, HomeSense, and Marshalls, you’ll find abundant opportunities to learn, thrive, and make an impact. Come join our TJX family—a Fortune 100 company and the world’s leading off-price retailer. Here at TJX Canada, we are an equal opportunity employer committed to the inclusion and accommodation of all individuals.

Job Description

We’re looking for a Site Reliability Engineer to help shape the future of intelligent operations by leveraging AIOps, GenAI, automation, and self‑healing technologies across enterprise platforms. In this role, you’ll drive initiatives that enhance system reliability, improve observability, and accelerate incident response while reducing manual effort through innovative automation solutions. You'll have the opportunity to influence modern engineering practices, work with cutting‑edge AI technologies, and make a measurable impact on the scalability and performance of critical business applications. Join a collaborative team where innovation, continuous improvement, and operational excellence are at the heart of everything we do.

Why Work With Us?
  • We value integrity, respect, and teamwork, encouraging a unique and inclusive culture.
  • Enjoy Associate discounts at our stores, available to you and eligible family members.
  • Immediate access to our Group Benefits package, including a Health Care Spending Account, Retirement Savings Program, Associate & Family Assistance Program, and various well‑being resources
  • A competitive vacation package, paired with a Vacation Trade Program that allows you to opt in for an extra week.
  • Comprehensive training and development resources designed to help you learn, grow, and succeed.
  • Exciting career paths with growth opportunities and tuition reimbursement to support your career progression.
What You’ll Do
  • Implement SRE best practices across production support, incident management, problem management, change management, release, and deployment processes.
  • Build and maintain automation scripts, runbooks, self‑healing workflows, and operational tools to reduce manual effort and improve MTTR.
  • Support AI‑driven operational capabilities such as incident summarization, alert enrichment, anomaly detection, log analysis, event correlation, and root cause recommendations.
  • Configure and enhance observability across logs, metrics, traces, dashboards, alerts, and synthetic monitoring.
  • Work with monitoring/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
  • Support implementation of SLIs, SLOs, SLAs, service health metrics, reliability KPIs, and operational dashboards.
  • Integrate monitoring, ITSM, CI/CD, cloud, and automation platforms using APIs, scripts, and workflows.
  • Participate in incident response, troubleshooting, RCA, postmortems, and corrective/preventive action planning.
  • Drive shift‑left reliability by embedding monitoring, alerting, automation, and operational readiness into SDLC and CI/CD pipelines.
  • Support reliable application and platform operations in Azure Cloud.
  • Create and maintain documentation, runbooks, knowledge articles, and automation playbooks.
  • Ensure AI‑enabled operations follow enterprise security, compliance, data privacy, and responsible‑AI guidelines.
About You
  • 6+ years of experience in SRE, DevOps, Production Support Engineering, Cloud Operations, Automation Engineering, or related roles.
  • Hands‑on experience in application support, incident management, problem management, change/release management, deployment support, monitoring, and documentation.
  • Hands‑on experience with observability/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
  • Experience with SLIs, SLOs, SLAs, operational KPIs, service health dashboards, and reliability metrics.
  • Strong coding/scripting experience in one or more languages such as Python, Java, PowerShell, Go, or Shell scripting.
  • Hands‑on experience with automation and DevOps tools such as Azure DevOps, GitHub Actions, Jenkins, Ansible, Terraform, Power Automate, Rundeck, or similar tools.
  • Knowledge of AIOps, Gen AI, or AI‑enabled operations use cases such as anomaly detection, log analysis, ticket classification, event correlation, and incident summarization.
  • Experience integrating enterprise tools using APIs, webhooks, scripts, and automation workflows.
  • Hands‑on experience supporting or implementing solutions in Azure Cloud.
  • Understanding of distributed systems, APIs, microservices, cloud‑native applications, and reliability engineering principles.
  • Strong troubleshooting, analytical, communication, and stakeholder management skills.
  • Ability to work across application, infrastructure, cloud, operations, security, and leadership teams.
Preferred Qualifications / Good To Have
  • Experience with Agentic AI, AI agents, or Gen AI‑based automation for IT operations.
  • Experience with Azure OpenAI, Microsoft Copilot Studio, LangChain, Semantic Kernel, vector databases, or RAG‑based solutions.
  • Experience with ITSM/incident response tools such as ServiceNow, Jira Service Management, PagerDuty, xMatters, Opsgenie, or similar platforms.
  • Experience with Docker, Kubernetes, OpenShift, CI/CD pipelines, Infrastructure as Code, GitOps, or DevSecOps.
  • Knowledge of machine learning concepts such as anomaly detection, classification, clustering, and time‑series analysis.
  • Understanding of Responsible AI, prompt engineering, model governance, data privacy, and security controls.
  • Experience in the Retail domain.
  • Relevant certifications in Azure, DevOps, SRE, AI, or ITIL.
Posting Details
  • Posting End Date: September 15, 2026
Additional Information

Candidates aged 18 and over will be required to undergo a criminal record check as part of the hiring process.

This job posting is for an existing position vacancy within our organization.

Address

60 Standish Court

Location

CAN Home Office Mississauga ON

Salary Range

Salary Range: $ 87,780.00-$122,892.00 /year *This represents the expected hiring range and may not represent the full pay range for the position. The salary offered may be higher than the posted range depending on several factors such as relevant skills, qualifications, and experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

The TJX Companies, Inc. • Mississauga

On-site
CAD 88,000 - 123,000
Employee discounts
Group benefits
Vacation package + extra week
+2
SRE Engineer
SRE Engineer

The TJX Companies, Inc. • Mississauga

On-site
CAD 88,000 - 123,000
Associate discounts at stores
Comprehensive group benefits package
Vacation trade program with extra week
+2
SRE Engineer
SRE Engineer

TJX Canada – Winners, Marshalls, HomeSense • Mississauga

On-site
CAD 88,000 - 123,000
Store discounts
Group Benefits package
Vacation & benefits program
+1
Senior Engineer
Senior Engineer

TJX Canada – Winners, Marshalls, HomeSense • Mississauga

On-site
CAD 88,000 - 123,000
Associate discounts
Group benefits (Health Care Spending,
Retirement Savings Program
+2
Winter Term (8 month) Business Intelligence & Automation Analyst Co-op
Winter Term (8 month) Business Intelligence & Automation Analyst Co-op

TJX Canada – Winners, Marshalls, HomeSense • Mississauga

Hybrid
CAD 62,000 - 86,000
Store discounts
Training & development
Assistance programs
+2
Staff Scrum Master
Staff Scrum Master

TJX Canada – Winners, Marshalls, HomeSense • Mississauga

On-site
CAD 105,000 - 155,000
Project Manager, Distribution Services - 12 Month Contract
Project Manager, Distribution Services - 12 Month Contract

The TJX Companies, Inc. • Mississauga

Hybrid
CAD 89,000 - 135,000
Associate discounts
Training & development
Group benefits
+1
Winter Term (4 month) Merchandise Operations Analyst Co-op/Intern
Winter Term (4 month) Merchandise Operations Analyst Co-op/Intern

TJX Canada – Winners, Marshalls, HomeSense • Mississauga

Hybrid
CAD 44,000 - 62,000
Employee discount
Hybrid work model
Training & development
Distribution Center Data Clerk
Distribution Center Data Clerk

TJX Canada – Winners, Marshalls, HomeSense • Mississauga

On-site
CAD 38,000 - 53,000
Associate discounts
Group Benefits
Vacation + Trade Program
+5
Project Manager, Distribution Services - 12 Month Contract
Project Manager, Distribution Services - 12 Month Contract

TJX Canada – Winners, Marshalls, HomeSense • Mississauga

Hybrid
CAD 89,000 - 135,000
Associate discounts
Comprehensive training and development
Group benefits package
+1