Core SRE Engineer

Acestack

Montreal

Hybrid

CAD 110,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Acestack is seeking a Core SRE Engineer in Montreal (hybrid) with 8+ years of IT experience to own production reliability across enterprise middleware and BI platforms. You will lead incident management, automate operations with Python, Shell scripting, Ansible and Terraform, and collaborate with engineering teams to reduce toil and improve service uptime.

This role requires strong Linux/Unix skills, monitoring with Splunk, Grafana, Prometheus, Loki, and a proactive approach to documentation and

Qualifications

  • 8+ years IT experience in SRE/production support.
  • Strong Linux/Unix administration.
  • Proficient in Python and Shell scripting.
  • Experience with Splunk, Grafana, Prometheus, Loki monitoring stacks.
  • Knowledge of Veritas Cluster Service, Load Balancers, VMware.
  • Understanding of ITIL and service management.

Responsibilities

  • Provide Level 3 SRE and production support for enterprise middleware and applications.
  • Manage and support tooling including Apache ZooKeeper, Ansible Automation Platform, and Terraform.
  • Act as the highest level of escalation for production incidents and service stability.
  • Troubleshoot complex Linux/Unix, application, middleware, VMware, and load-balancer issues.
  • Monitor health using Splunk, Grafana, Prometheus, Loki.
  • Collaborate with engineering to resolve issues and improve reliability.
  • Automate operational processes with Python, Shell, Ansible, and Terraform.
  • Support code releases and coordinate deployments.
  • Manage production escalations and participate in on-call/weekend support.
  • Prepare operational reports and coordinate with stakeholders.
  • Maintain documentation and knowledge transfer across global teams.

Skills

SRE
Linux/Unix
Python
Shell scripting
Automation
Monitoring stacks
Incident management
VMware
ITIL
BI platforms

Tools

Ansible Automation Platform
Terraform
Apache ZooKeeper
Load Balancers
Splunk
Grafana
Prometheus
Loki

Job description

Job Title: Core SRE Engineer
Job Type: Full Time
Location: Montreal, QC (Hybrid)

Job Description

We are seeking an experienced Core SRE Engineer with 8+ years of IT experience and strong expertise in Linux/Unix, Python, Shell Scripting, monitoring, application support, and site reliability engineering. The ideal candidate will have hands-on experience supporting enterprise middleware and BI platforms, managing production environments, troubleshooting complex incidents, and improving system reliability through automation.

The role involves L3/global production support, infrastructure management, incident and problem management, automation, and close collaboration with engineering and development teams.

Key Responsibilities
  • Provide Level 3 SRE and production support for enterprise middleware and application platforms.
  • Manage and support applications and tooling involving Apache ZooKeeper, Ansible Automation Platform, and Terraform.
  • Act as the highest level of escalation for production incidents and service stability issues.
  • Troubleshoot complex Linux/Unix, application, middleware, VMware, and load-balancer issues.
  • Monitor application and infrastructure health using Splunk, Grafana, Prometheus, and Loki.
  • Participate in incident, change, escalation, and problem management activities.
  • Collaborate with engineering and development teams to resolve production issues and improve service reliability.
  • Automate operational processes using Python, Shell scripting, Ansible, and Terraform to reduce manual effort and operational toil.
  • Support code releases and coordinate with development teams during application deployments.
  • Manage production escalations and participate in on-call/weekend support as required.
  • Prepare and submit operational reports and coordinate with multiple stakeholders.
  • Maintain strong documentation and ensure effective knowledge transfer across global teams.
Required Skills & Qualifications
  • 8+ years of overall IT experience, with 8+ years in SRE or a similar production support role.
  • Advanced hands-on experience with Linux/Unix administration and support.
  • Strong Shell scripting and Python programming skills.
  • Experience with Splunk and/or Grafana, Prometheus, and Loki monitoring stacks.
  • Working knowledge of Veritas Cluster Service, Load Balancers, and VMware.
  • Strong understanding of ITIL principles and IT service management practices.
  • Experience supporting BI platforms and enterprise middleware environments.
  • Strong troubleshooting and outage-management capabilities.
  • Excellent written and verbal communication skills.
Preferred Skills
  • Experience with Ansible playbooks and Ansible Automation Platform administration.
  • Experience with Terraform, particularly Terraform Enterprise.
  • Knowledge of Docker and Kubernetes/OpenShift.
  • Experience with Git, Bitbucket, and CI/CD toolchains.
  • Knowledge of Agile methodologies.
  • Good understanding of JVMs and garbage collection mechanisms.
  • Experience with relational databases.
  • Application support, production release, and development team coordination experience.
Key Competencies
  • Strong analytical and problem-solving skills.
  • Ability to manage multiple priorities in high-pressure production environments.
  • Strong ownership and escalation-management capabilities.
  • Ability to automate repetitive operational processes and reduce system downtime.
  • Excellent stakeholder coordination and communication skills.
  • Willingness to participate in weekend and on-call support.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Level 3 Support and SRE
Level 3 Support and SRE

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 75,000 - 95,000
Senior Core SRE Engineer | Automation & Reliability
Senior Core SRE Engineer | Automation & Reliability

Acestack • Montreal

Hybrid
CAD 110,000 - 150,000
Platform & SRE Engineer
Platform & SRE Engineer

Aarorn Technologies Inc • Montreal (administrative region)

On-site
CAD 5,786,000 - 7,990,000
SRE x 2
SRE x 2

HRB • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
SRE Engineer
SRE Engineer

SFE • Toronto

On-site
CAD 90,000 - 130,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Production Management Lead
Production Management Lead

Compunnel, Inc. • Montreal (administrative region)

On-site
CAD 85,000 - 110,000
Opportunities for personal development
Career growth potential
MONTREAL [Hybrid] - Senior DevOps SRE
MONTREAL [Hybrid] - Senior DevOps SRE

QUANTEAM (RAINBOW PARTNERS Group) • Montreal

On-site
CAD 90,000 - 130,000
Hybrid work model
Site Reliability Engineer (Linux / Cloud Infrastructure)
Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT Group • Montreal

On-site
CAD 80,000 - 100,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 130,000