Site Reliability Engineer

Mantu

Montreal (administrative region)

On-site

CAD 90,000 - 130,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mantu is seeking a hands-on SRE / Platform Engineer to operate and continuously improve a Kubernetes-based platform used across the organization. You will focus on day‑to‑day operations, security, performance, and automation to support a diverse technology stack in cloud and on‑prem environments.

You will work on incident response, observability enhancements, and tooling to improve developer experience while expanding platform capabilities and reliability.

Qualifications

  • 5-7 years of experience in SRE, DevOps, Platform Engineering, or similar infrastructure-focused role.
  • Hands-on experience operating Kubernetes in production, ideally with Service Mesh exposure.
  • Excellent Linux and command-line skills with the ability to troubleshoot complex systems.
  • Experience with a major public cloud platform, preferably Azure or AWS.
  • Familiarity with Grafana, Prometheus, Loki, Tempo; Python or Java scripting is an asset.
  • Knowledge of CI/CD and Infrastructure as Code, particularly Helm or Terraform.

Responsibilities

  • Operate and maintain a Kubernetes-based platform across public, private, and on-prem environments.
  • Monitor platform health using observability tools, ensuring actionable alerts and documented runbooks.
  • Manage incidents end-to-end including incident response, troubleshooting, and post-incident reviews.
  • Build automation, diagnostic tools, and performance tests to improve reliability.
  • Support client onboarding, platform upgrades, infrastructure changes, and capacity management.
  • Identify and implement continuous improvements across operations and deployment processes.

Skills

Kubernetes in production
Linux
Python/Java scripting
CI/CD
Service Mesh

Tools

Grafana
Prometheus
Loki
Tempo
Helm
Terraform

Job description

Join an API Platform team responsible for operating and continuously improving a critical enterprise platform used by development teams across the organization. The team follows an API-first approach and supports a diverse technology landscape spanning public cloud, hybrid cloud, and on-premises environments.

This is a hands‑on SRE / Platform Engineering opportunity that combines software engineering and operations. You will help keep the Kubernetes platform stable, secure, performant, and scalable while building automation, improving observability, supporting production incidents, and enhancing the overall developer experience. The role is focused on the day‑to‑day operation and continuous improvement of the platform, rather than traditional application development.

Main Responsibilities
  • Operate and maintain a Kubernetes-based platform across public cloud, private cloud, and on-premises environments.
  • Monitor platform health using observability tools, ensuring alerts are actionable and supported by accurate runbooks and documentation.
  • Manage incidents end-to-end, including incident response, troubleshooting, root cause analysis, and post-incident reviews.
  • Build automation, diagnostic tools, and performance tests to reduce manual effort and improve platform reliability.
  • Support client onboarding, platform upgrades, infrastructure changes, and capacity management to ensure the platform scales with demand.
  • Identify and implement continuous improvements across operations, automation, deployment, monitoring, and support processes.
Qualifications
  • 5-7 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or a similar infrastructure-focused role.
  • Strong hands‑on experience operating Kubernetes in production, ideally with Service Mesh exposure.
  • Excellent Linux and command‑line skills with the ability to troubleshoot complex systems from application to infrastructure layers.
  • Experience with a major public cloud platform, preferably Azure or AWS.
  • Familiarity with Grafana, Prometheus, Loki, and Tempo; Python or Java scripting experience is an asset.
  • Knowledge of CI/CD and Infrastructure as Code, particularly Helm or Terraform, combined with strong analytical and problem‑solving skills.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Open Systems Technologies • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind Americas • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind • Montreal (administrative region)

Hybrid
CAD 120,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
Site Reliability Engineer (SRE) – UI/UX
Site Reliability Engineer (SRE) – UI/UX

Software Mind • Montreal (administrative region)

Hybrid
CAD 110,000 - 165,000
Competitive salary
Laptop provided
Professional development
+2
Senior SRE: Global SaaS Platform, Kubernetes & Cloud
Senior SRE: Global SaaS Platform, Kubernetes & Cloud

Kong Inc. • Toronto

On-site
CAD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000
Senior SRE – Kubernetes for Production UI/AI Stack
Senior SRE – Kubernetes for Production UI/AI Stack

Software Mind • Montreal (administrative region)

Hybrid
CAD 110,000 - 165,000
Competitive salary
Laptop provided
Professional development
+2
Kubernetes Platform SRE: Incidents, Automation & Observability
Kubernetes Platform SRE: Incidents, Automation & Observability

Open Systems Technologies • Montreal (administrative region)

On-site
CAD 90,000 - 130,000