A leading financial services firm in Montreal is seeking an experienced Site Reliability Engineer to enhance operational resilience across cloud platforms and on-prem environments. The role involves incident management, advanced observability, and implementation of auto-healing tooling. Ideal candidates will have over eight years of relevant experience, solid software engineering skills, and be bilingual in French and English. The position offers a competitive salary, flexible work arrangements, and the potential for bonuses based on performance.
Qualifications
8+ years of experience in SRE/Platform/Infrastructure/Software Engineering.
Solid software engineering skills in at least one of: Go, Python, or TypeScript.
Bilingual (French and English) to interact with clientele.
Responsibilities
Lead high-severity investigations and RCA with App/Infra/Incident teams.
Implement end-to-end traces/metrics/logs with consistent semantics.
Build policy-driven remediation and enable progressive delivery.
Skills
OpenTelemetry instrumentation
Dynatrace
Kubernetes
Python
Go
Terraform
GitOps
Chaos engineering
Job description
A leading financial services firm in Montreal is seeking an experienced Site Reliability Engineer to enhance operational resilience across cloud platforms and on-prem environments. The role involves incident management, advanced observability, and implementation of auto-healing tooling. Ideal candidates will have over eight years of relevant experience, solid software engineering skills, and be bilingual in French and English. The position offers a competitive salary, flexible work arrangements, and the potential for bonuses based on performance.