Site Reliability Specialist, IT Operations

Sherweb Inc.

Quebec

Hybrid

CAD 71,000 - 102,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Beautiful offices
Cutting-edge tools
Paid time off starting day one
Employee benefits

Job summary

Sherweb Inc. in Québec is seeking a Site Reliability Specialist to join the IT Operations team.

The role combines systems administration, software development, automation, observability, and SRE practices to improve service reliability and scalability. Working closely with Infrastructure, Development, DevOps, Platform, Security, and Product teams, you will implement reliability standards, monitor performance, and drive automation to reduce toil while supporting production services in a hybrid

Qualifications

  • College or university degree in computer science, software development, information technology, engineering, or a combination of equivalent training and experience.

Responsibilities

  • Develop, maintain, and improve scripts, automation workflows, and operational tooling to reduce manual intervention, improve reliability, and lower operational toil.
  • Apply SRE principles to improve the reliability, availability, performance, and resilience of Sherweb’s hosted platforms and production services.
  • Implement and support reliability standards, service level objectives (SLOs), service level indicators (SLIs), and operational practices established for platforms and services.
  • Use a developer mindset to transform repetitive operational tasks into scalable, reusable, documented, and supportable automation.
  • Provide advanced operational support and resolve incidents affecting production services while collaborating closely with SRE, Infrastructure, Development, DevOps, and Product teams.

Skills

PowerShell
Python
Bash
JavaScript
TypeScript
C#
Windows Server
Linux
Observability
Git
CI/CD
Automation

Education

College or university degree in computer science, software development, information technology, engineering

Tools

Terraform
Ansible
DSC
Azure DevOps
GitHub Actions
Docker
Kubernetes
Azure AI Foundry
Power Automate
Copilot/AI agents

Job description

Job Details
Posting Details
  • Posted on July 28, 2026
Locations
  • Showing 1 location
  • Québec, Province
  • Hybrid
  • .IT Operations & Delivery
  • Full-Time
  • Requisition #: SITER002792
Description

Location : Hybrid - Province of Quebec

The Site Reliability Specialist on the IT Operations team contributes to the reliability, availability, performance, and resilience of Sherweb's platforms and services.

This is a highly technical individual contributor role that applies Site Reliability Engineering (SRE) principles to production environments. The role combines systems administration, software development, automation, observability, and operational excellence to improve service reliability, reduce operational toil, and increase platform scalability.

Working closely with Infrastructure, Development, DevOps, Platform, Security, and Product teams, the Site Reliability Specialist helps ensure production systems remain stable, supportable, and continuously improving through engineering and automation practices.

Here's how you will contribute to the success of the company

  • Develop, maintain, and improve scripts, automation workflows, and operational tooling to reduce manual intervention, improve reliability, and lower operational toil.
  • Apply SRE principles to improve the reliability, availability, performance, and resilience of Sherweb’s hosted platforms and production services.
  • Implement and support reliability standards, service level objectives (SLOs), service level indicators (SLIs), and operational practices established for platforms and services.
  • Use a developer mindset to transform repetitive operational tasks into scalable, reusable, documented, and supportable automation.
  • Provide advanced operational support and resolve incidents affecting production services while collaborating closely with SRE, Infrastructure, Development, DevOps, and Product teams.
  • Investigate recurring issues and perform root cause analysis to identify short-term corrective actions and long-term reliability improvements.
  • Build, support, and maintain production systems and hosted service technologies while following operational procedures, security best practices, and compliance requirements.
  • Improve monitoring, alerting, logging, metrics, and operational visibility to help the team detect issues earlier, understand system behavior, and prevent incidents.
  • Contribute to improving end-to-end observability and system understanding through metrics, logs, traces, telemetry, and operational diagnostics.
  • Contribute to observability-as-code, infrastructure-as-code, configuration-as-code, and automation practices where applicable.
  • Explore and leverage Azure AI Foundry, Power Automate, and AI agent capabilities to improve operational efficiency, automate repetitive workflows, and accelerate incident response or service reliability improvements.
  • Participate in platform lifecycle activities, deployments, maintenance windows, migrations, and continuous service improvement initiatives.
  • Collaborate with developers, architects, subject matter experts, DevOps, and infrastructure teams to support implementation, optimization, troubleshooting, and operational readiness of services.
  • Create and maintain operational documentation, including SOPs, runbooks, troubleshooting guides, automation documentation, maintenance procedures, and knowledge-sharing materials.
  • Track, organize, and manage incidents, requests, and service tickets to respect SLAs and ensure clear communication through the ITSM process.
  • Participate in rotational on-call duty and perform maintenance work outside normal business hours when required.
  • Carry out all other related tasks per the job’s evolution and departmental needs.

Here's what you need to have and master to get the job

Education

  • College or university degree in computer science, software development, information technology, engineering, or a combination of equivalent training and experience.

Experience

  • 3 to 5 years of experience in systems administration, IT operations, infrastructure support, software development, DevOps, automation, or a similar technical role.
  • Experience supporting production systems in business-critical and customer-facing environments.
  • Proven experience improving operational efficiency through automation and engineering practices.

Core Skills

  • Strong scripting or development skills with at least one language such as PowerShell, Python, Bash, JavaScript, TypeScript, or C#.
  • Proven ability to design, write, test, troubleshoot, document, and maintain scripts or small applications used to automate operational tasks.
  • Proven experience supporting Microsoft and/or Linux server environments, including troubleshooting, maintenance, operational support, and automation.
  • Strong diagnostic, investigation, and problem-solving skills with the ability to analyze incidents, identify root causes, and implement sustainable improvements through automation or engineering practices.
  • Good understanding of distributed systems, networking concepts, system dependencies, availability, performance, reliability, and service operations in production environments.
  • Experience with monitoring, alerting, observability, log management, telemetry and operational data analysis to detect, troubleshoot, and prevent issues.
  • Familiarity with version control, Git-based workflows, CI/CD pipelines, code review practices, and deployment automation.
  • Experience with infrastructure as code, configuration management, or automation tools such as Terraform, Ansible, DSC, Azure DevOps, GitHub Actions, Docker, or Kubernetes is an asset.
  • Familiarity with Azure AI Foundry, Power Automate, Copilot/AI agents, or agent-based automation concepts is an asset.
  • Knowledge of high availability environments, virtualization, cloud services, backup and restore practices, and production support models is an asset.

Professional Attributes

  • Autonomous, reliable, and motivated, with a continuous learning mindset and a strong interest in improving reliability through software engineering and automation practices.
  • Strong communication and collaboration skills, with the ability to work effectively with technical and non-technical stakeholders.
  • Excellent English skills, both spoken and written, are essential; fluency in French is an asset.
  • Relevant industry certifications such as Microsoft Azure, Red Hat, Linux Foundation, Kubernetes, DevOps, or observability platforms are considered assets.

Additional Requirements

  • Availability for rotation on the on-call schedule in a 24/7 environment.

Benefits of working at Sherweb

Sherweb is first and foremost a culture where our customers’ needs are at the heart of everything we do, supported by Sherwebers committed to living our values of passion, teamwork, and integrity.

Dynamic and flexible work environment

  • Beautiful offices designed for collaboration (Sherbrooke and downtown Montreal)
  • Access tocutting-edgetools and technologies
  • A results-driven culture where ideas, actions, andexpertiseare recognized
  • Supportive colleagues from diverse professional and cultural backgrounds
  • A generous employee referral program

Flexible total compensation package

  • Base salary between $71,330.00 and $101,900.00 per year
  • Vacation allowance based on prior experience
  • Paid time off starting on day one (vacation and personal days)
  • Employee benefits program with savings at local merchants

Unlimited growth opportunities

  • Close relationship with your direct manager and open communication to support your development
  • Multiple learning and development opportunities
  • Tools to track your progress and support career growth

A vibrant social life (Sherweblife)

A rich calendar of virtual and in-person activities designed to foster connection throughout the year.

At Sherweb, we believe in transparency and pay equity. The salary range provided is intended to give an indication of what you can expect for this role. However, we recognize that each candidate brings a unique set of skills and experiences. The final compensation package will be tailored to reflect the selected candidate’s qualifications and expertise, ensuring we remain competitive and fair in our offers.

Reasons for the requirement of English: Sherweb has international customers and fluency in English is the only way to ensure proper service delivery to them. The main tasks related to this position require written and oral communication with an English-speaking clientele at all times.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Specialist, IT Operations
Site Reliability Specialist, IT Operations

Sherweb • Quebec

On-site
CAD 71,330 - 101,900
Hybrid work model
Office locations in Sherbrooke and MTL
Salary transparency
System Administrator, IT Operations
System Administrator, IT Operations

Sherweb Inc. • Montreal (administrative region), Sherbrooke

Hybrid
CAD 62,000 - 89,000
Salary range 62k-89k CAD
Vacation time
Advanced paid hours
+1
Senior Business Analyst
Senior Business Analyst

Sherweb • Quebec

Hybrid
CAD 81,000 - 116,000
Flexible benefits plan
Flexible savings fund option
Monthly home internet allowance
+1
Technical Advisor II
Technical Advisor II

Sherweb Inc. • Sherbrooke

On-site
CAD 37,000 - 53,000
Annual salary review
Flexible benefits plan
Home internet allowance
+2
Business Analyst, Technical
Business Analyst, Technical

Sherweb • Sherbrooke

Hybrid
CAD 81,000 - 116,000
Annual salary review
Flexible benefits
Home internet allowance
Product Manager, Business Apps & AI
Product Manager, Business Apps & AI

Sherweb • Quebec

On-site
CAD 81,000 - 116,000
Annual salary review based on perfomrm
Flexible benefits plan
Home internet allowance
+1
Remote Technical Fellow, Azure
Remote Technical Fellow, Azure

Bilinguallink • Surrey

Remote
CAD 145,000 - 208,000
Monthly home internet allowance
Flexible benefits plan
Vacation time based on experience
+1
Technical Fellow, Microsoft Copilot & AI
Technical Fellow, Microsoft Copilot & AI

Sherweb • Canada

On-site
CAD 142,000 - 202,000
Home internet allowance
Flexible benefits plan
Vacation based on experience
+2
Growth & Content Marketing Specialist (Bilingual)
Growth & Content Marketing Specialist (Bilingual)

Sherpa • Montreal (administrative region)

On-site
CAD 70,000 - 90,000
Hybrid work environment
Competitive compensation
Global projects
CONSEILLER.ERE EN ANALYSE STRATÉGIQUE ET EXPLOITATION DE DONNÉES
CONSEILLER.ERE EN ANALYSE STRATÉGIQUE ET EXPLOITATION DE DONNÉES

Go RH • Sherbrooke

On-site
CAD 34,440 - 44,083
Environnement de travail stimulant
Équilibre travail-vie personnelle
Opportunités de développement professionnel