Core SRE Engineer

REALIGN LLC

Montreal (administrative region)

Hybrid

CAD 110,000 - 160,000

Full time

30 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

REALIGN LLC in Montreal, QC is seeking a highly experienced Core SRE Engineer (8+ years) to own production reliability across enterprise middleware and BI platforms. You will tackle incident management, automation, and on-call support, with emphasis on Linux/Unix, Python, Shell scripting, and monitoring stacks.

You will collaborate with engineering teams to reduce downtime, improve service quality, and implement scalable infrastructure through Terraform, Ansible, and VMware in a hybrid work

Qualifications

  • 8+ years IT experience in SRE or production support.
  • Advanced Linux/Unix administration and support.
  • Experience with monitoring stacks: Splunk, Grafana, Prometheus, Loki.

Responsibilities

  • Provide Level 3 SRE and production support for enterprise middleware and platforms.
  • Manage and support applications with Ansible, Terraform, and ZooKeeper.
  • Act as highest escalation for production incidents and stability issues.
  • Troubleshoot Linux/Unix, middleware, VMware, and load-balancer problems.
  • Monitor health with Splunk, Grafana, Prometheus, Loki.
  • Automate operations using Python, Shell, Ansible, Terraform.

Skills

Linux/Unix
Python
Shell scripting
Monitoring
ITIL
Troubleshooting
SRE/production support
Communication
On-call support
Documentation

Tools

Apache ZooKeeper
Ansible Automation Platform
Terraform
Splunk
Grafana
Prometheus
Loki
Veritas Cluster Service
VMware
Load Balancers

Job description

Job Title: Core SRE Engineer
Job Type: Full Time
Location: Montreal, QC (Hybrid)

Job Description

We are seeking an experienced Core SRE Engineer with 8+ years of IT experience and strong expertise in Linux/Unix, Python, Shell Scripting, monitoring, application support, and site reliability engineering. The ideal candidate will have hands‑on experience supporting enterprise middleware and BI platforms, managing production environments, troubleshooting complex incidents, and improving system reliability through automation.

The role involves L3/global production support, infrastructure management, incident and problem management, automation, and close collaboration with engineering and development teams.

Key Responsibilities
  • Provide Level 3 SRE and production support for enterprise middleware and application platforms.
  • Manage and support applications and tooling involving Apache ZooKeeper, Ansible Automation Platform, and Terraform.
  • Act as the highest level of escalation for production incidents and service stability issues.
  • Troubleshoot complex Linux/Unix, application, middleware, VMware, and load‑balancer issues.
  • Monitor application and infrastructure health using Splunk, Grafana, Prometheus, and Loki.
  • Participate in incident, change, escalation, and problem management activities.
  • Collaborate with engineering and development teams to resolve production issues and improve service reliability.
  • Automate operational processes using Python, Shell scripting, Ansible, and Terraform to reduce manual effort and operational toil.
  • Support code releases and coordinate with development teams during application deployments.
  • Manage production escalations and participate in on‑call/weekend support as required.
  • Prepare and submit operational reports and coordinate with multiple stakeholders.
  • Maintain strong documentation and ensure effective knowledge transfer across global teams.
Required Skills & Qualifications
  • 8+ years of overall IT experience, with 8+ years in SRE or a similar production support role.
  • Advanced hands‑on experience with Linux/Unix administration and support.
  • Strong Shell scripting and Python programming skills.
  • Experience with Splunk and/or Grafana, Prometheus, and Loki monitoring stacks.
  • Working knowledge of Veritas Cluster Service, Load Balancers, and VMware.
  • Strong understanding of ITIL principles and IT service management practices.
  • Experience supporting BI platforms and enterprise middleware environments.
  • Strong troubleshooting and outage‑management capabilities.
  • Excellent written and verbal communication skills.
Preferred Skills
  • Experience with Ansible playbooks and Ansible Automation Platform administration.
  • Experience with Terraform, particularly Terraform Enterprise.
  • Knowledge of Docker and Kubernetes/OpenShift.
  • Experience with Git, Bitbucket, and CI/CD toolchains.
  • Knowledge of Agile methodologies.
  • Good understanding of JVMs and garbage collection mechanisms.
  • Experience with relational databases.
  • Application support, production release, and development team coordination experience.
  • Strong analytical and problem‑solving skills.
  • Ability to manage multiple priorities in high‑pressure production environments.
  • Strong ownership and escalation‑management capabilities.
  • Ability to automate repetitive operational processes and reduce system downtime.
  • Excellent stakeholder coordination and communication skills.
  • Willingness to participate in weekend and on‑call support.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Level 3 Support and SRE
Level 3 Support and SRE

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 75,000 - 95,000
Core SRE Engineer - Production Reliability & Automation
Core SRE Engineer - Production Reliability & Automation

REALIGN LLC • Montreal (administrative region)

Hybrid
CAD 110,000 - 160,000
Administrateur Système Linux (Senior)
Administrateur Système Linux (Senior)

SII Canada • Montreal (administrative region)

On-site
CAD 70,000 - 90,000
SRE x 2
SRE x 2

HRB • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Production Support
Production Support

Axelon Services Corporation • Montreal (administrative region)

Hybrid
CAD 75,000 - 95,000
Production Management Lead
Production Management Lead

Compunnel, Inc. • Montreal (administrative region)

On-site
CAD 85,000 - 110,000
Opportunities for personal development
Career growth potential
MONTREAL [Hybrid] - Senior DevOps SRE
MONTREAL [Hybrid] - Senior DevOps SRE

QUANTEAM (RAINBOW PARTNERS Group) • Montreal

Hybrid
CAD 90,000 - 130,000
Hybrid work model
Montreal [hybrid] Front Office SRE Support Analyst
Montreal [hybrid] Front Office SRE Support Analyst

QUANTEAM - North America (RAINBOW PARTNERS Group) • Montreal (administrative region)

On-site
CAD 90,000 - 140,000