Site Reliability Engineer

ITC Infotech

Vancouver

On-site

CAD 140,000 - 180,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ITC Infotech is seeking an experienced Sr DevOps and Site Reliability Engineer to maintain and scale AWS, PostgreSQL, Airflow and Snowflake environments. You will ensure data governance and reliability across the Retail Planning data ecosystem while partnering with Planning and Analytics teams.

The role emphasizes IaC, CI/CD, automation, and incident management to improve delivery efficiency and platform resilience. Vancouver-based team, on-site collaboration, and growth opportunities.

Qualifications

  • 8+ years of experience in DevOps, Cloud Engineering, Platform Engineering, or SRE roles.
  • Strong hands-on experience with AWS cloud services.
  • Expertise in GitLab CI/CD pipeline development and deployment automation.
  • Extensive experience with Terraform for infrastructure provisioning and management.
  • Strong Python scripting and automation skills.
  • Deep understanding of Kubernetes architecture, administration, and troubleshooting.
  • Hands-on experience with Helm chart development and deployment.
  • Experience implementing GitOps workflows with ArgoCD.
  • Strong experience with monitoring, logging, and observability platforms like Datadog and Splunk.

Responsibilities

  • Design, implement, and maintain scalable cloud infrastructure on AWS.
  • Build and manage Infrastructure as Code (IaC) using Terraform.
  • Develop and maintain CI/CD pipelines using GitLab CI/CD.
  • Automate operational processes and workflow integrations using Python.
  • Manage containerized environments using Kubernetes, Helm, and Rancher.
  • Establish infrastructure standards, reusable modules, and deployment best practices.
  • Collaborate with development, architecture, and security teams to improve software delivery efficiency.
  • Ensure platform availability, performance, scalability, and reliability.
  • Define and manage SLIs, SLOs, and SLAs.
  • Lead incident management, root cause analysis, and post-incident reviews.
  • Drive automation initiatives to reduce operational overhead and improve resilience.
  • Participate in capacity planning and performance optimization.

Skills

AWS
GitLab CI/CD
Terraform
Python
Kubernetes
ArgoCD
Linux
Networking
Observability

Education

Bachelor / Master in CS, IT, Engineering

Tools

Rancher
Helm
Datadog
Splunk

Job description

ITC Infotech is a leading global technology services and solutions provider, led by Business and Technology Consulting. ITC Infotech provides business-friendly solutions to help clients succeed and be future-ready, by seamlessly bringing together digital expertise, strong industry specific alliances and the unique ability to leverage deep domain expertise from ITC Group businesses. We provide technology solutions and services to enterprises across industries such as Banking & Financial Services, Healthcare, Manufacturing, Consumer Goods, Travel and Hospitality, through a combination of traditional and newer business models, as a long-term sustainable partner.

Role Summary

We’re looking for an experienced Sr DevOps and Site Reliability Engineer to maintain and scale AWS, PostgreSQL (ODS and transactional), Airflow and Snowflake environments. Should also be comfortable with replication, failover mechanisms and backup/recovery verifications mainly in the Snowflake space. Understanding various models and curated data sets that power Retail Planning analytics and decision-making (e.g., demand planning, assortment, allocation, replenishment, merchandise/financial planning, and performance reporting). You’ll partner closely with Planning stakeholders, Analytics/BI, and upstream source teams to ensure high-quality, governed, and reliable data products across the Retail Planning data ecosystem.

Key Responsibilities
DevOps & SRE Role
  • Design, implement, and maintain scalable cloud infrastructure on AWS.
  • Build and manage Infrastructure as Code (IaC) using Terraform.
  • Develop and maintain CI/CD pipelines using GitLab CI/CD.
  • Automate operational processes and workflow integrations using Python.
  • Manage containerized environments using Kubernetes, Helm, and Rancher.
  • Establish infrastructure standards, reusable modules, and deployment best practices.
  • Collaborate with development, architecture, and security teams to improve software delivery efficiency.
  • Ensure platform availability, performance, scalability, and reliability.
  • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and SLAs.
  • Implement observability and monitoring strategies using Datadog and Splunk.
  • Lead incident management, root cause analysis, and post-incident reviews.
  • Drive automation initiatives to reduce operational overhead and improve system resilience.
  • Conduct capacity planning, performance optimization, and reliability assessments.
  • Participate in on-call rotations and support critical production environments.
Stakeholder Collaboration
  • Work closely with Retail Planning, Merchandising, and Supply Chain stakeholders to evaluate upcoming implementations and relevant infrastructure impacts.
  • Participate and contribute to BI/Analytics teams to automate recurring data preparation, KPI calculations, and reporting datasets.
  • Participate and contribute in incident triage and root-cause analysis for platform related pipeline/data issues, drive prevention through durable fixes.
Required Skills & Qualifications
  • 8+ years of experience in DevOps, Cloud Engineering, Platform Engineering, or SRE roles.
  • Strong hands‑on experience with AWS cloud services.
  • Expertise in GitLab CI/CD pipeline development and deployment automation.
  • Extensive experience with Terraform for infrastructure provisioning and management.
  • Strong Python scripting and automation skills.
  • Deep understanding of Kubernetes architecture, administration, and troubleshooting.
  • Hands‑on experience with Helm chart development and deployment.
  • Experience implementing GitOps workflows with ArgoCD.
  • Strong experience with monitoring, logging, and observability platforms:
  • Strong Linux administration and troubleshooting skills.
  • Experience with networking concepts including DNS, Load Balancers, VPCs, Security Groups, and IAM.
  • Familiarity with retail planning concepts and data: forecasting, inventory, allocation, replenishment, assortment, pricing/promo, sell-through, WOS/DOH, in-season vs. pre-season planning, etc.
  • Exposure to common retail data sources like sales transactions, inventory snapshots, product hierarchies, store/channel attributes, supplier/lead times, and order flows.
Education Qualification
  • Bachelor / Master in Computer Science, Information Technology, Engineering, or related field (equivalent experience acceptable).

ITC Infotech is an Equal Opportunity Employer. We believe that no one should be discriminated against because of their differences, such as age, disability, ethnicity, gender, gender identity and expression, religion, or sexual orientation. All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by federal, state, or local law. ITC infotech is committed to providing veteran employment opportunities to our service men and women.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud SRE & DevOps Engineer – AWS, Kubernetes, CI/CD
Senior Cloud SRE & DevOps Engineer – AWS, Kubernetes, CI/CD

ITC Infotech • Vancouver

On-site
CAD 140,000 - 180,000
Lead Data Engineer
Lead Data Engineer

ITC Infotech • Vancouver

On-site
CAD 150,000 - 190,000
Technical Delivery Manager
Technical Delivery Manager

ITC Infotech • Vancouver

Hybrid
CAD 120,000 - 150,000
Senior Program Manager
Senior Program Manager

ITC Infotech • Vancouver

On-site
CAD 120,000 - 150,000
Senior DevOps Engineer
Senior DevOps Engineer

Zoom Information, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior Site Reliability Developer
Senior Site Reliability Developer

United States Digital Space LLC • Toronto

On-site
CAD 107,000 - 157,000
Salary transparency
In-person onboarding
Site Reliability Engineer
Site Reliability Engineer

RXinsider LTD. • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

Hybrid
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4
Manager, Site Reliability Operations
Manager, Site Reliability Operations

Canadian Tire Corporation • Toronto

On-site
CAD 79,000 - 115,000
Comprehensive benefits and retirement programs
Performance incentives
Mental health benefits up to $5,000
+3
Site Reliability / DevOps Engineer
Site Reliability / DevOps Engineer

Infotek Consulting Inc. • Toronto

Hybrid