Site Reliability Engineer

Open Systems Technologies

Montreal (administrative region)

On-site

CAD 90,000 - 130,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Open Systems Technologies in Montreal is seeking an experienced DevOps professional to operate and maintain a Kubernetes-based platform across public and private clouds. You will manage incidents end-to-end, onboard new clients, and build automation to improve performance and reduce manual effort.

The role requires hands-on Kubernetes in production, strong Linux skills, and experience with cloud providers such as Azure or AWS.

Qualifications

  • Hands on experience operating Kubernetes in production, ideally with Service Mesh.
  • Strong Linux and command line fundamentals.
  • Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure.
  • Experience working with a public cloud provider, preferable Azure or AWS.
  • Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus.
  • Scripting or coding in Python or Java is a strong plus.
  • CI/CD, infrastructure as code such as Helm or Terraform is a plus
  • A financial services background is not required.

Responsibilities

  • Operate and maintain the Kubernetes based platform across public and private cloud environments.
  • Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation.
  • Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews.
  • Onboard and support new clients onto the platform.
  • Build automation and diagnostic tooling that cuts manual effort and evaluates performance.
  • Identify and deliver process improvements.
  • Support software and hardware upgrades and keep components up to date.
  • Proactively manage capacity so the platform scales with demand.

Skills

Kubernetes in production
Linux fundamentals
Debugging and troubleshooting
Public cloud experience (Azure/AWS)
CI/CD
Python
Java

Tools

Grafana
Prometheus
Loki
Tempo
Python
Java
Helm
Terraform

Job description

  • Operate and maintain the Kubernetes based platform across public and private cloud environments.
  • Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation.
  • Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews.
  • Onboard and support new clients onto the platform.
  • Build automation and diagnostic tooling that cuts manual effort and evaluates performance.
  • Identify and deliver process improvements.
  • Support software and hardware upgrades and keep components up to date.
  • Proactively manage capacity so the platform scales with demand.

Required Skills:

  • Hands on experience operating Kubernetes in production, ideally with Service Mesh.
  • Strong Linux and command line fundamentals.
  • Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure.
  • Experience working with a public cloud provider, preferable Azure or AWS.
  • Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus.
  • Scripting or coding in Python or Java is a strong plus.
  • CI/CD, infrastructure as code such as Helm or Terraform is a plus
  • A financial services background is not required.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Mantu • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind Americas • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Kubernetes Platform SRE: Incidents, Automation & Observability
Kubernetes Platform SRE: Incidents, Automation & Observability

Open Systems Technologies • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
DevOps Specialist - Kubernetes
DevOps Specialist - Kubernetes

Myticas Consulting • Ottawa

On-site
CAD 90,000 - 130,000
Senior DevOps Engineer
Senior DevOps Engineer

Mphasis • Toronto

On-site
CAD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

MarkiTech • Toronto

On-site
CAD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Vertex Elite LLC • Ottawa

On-site
CAD 83,000 - 124,000
Platform Engineer
Platform Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 80,000 - 120,000