[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Worky

Montreal (administrative region)

On-site

CAD 120,000 - 170,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Laptop
Flexible work arrangements
Professional development and training

Job summary

Software Mind is seeking a Senior SRE to support production reliability for a Kubernetes-based UI service and AI framework. The role emphasizes Kubernetes operations, observability, and incident ownership across distributed systems.

Ideal candidates bring hands-on Kubernetes, Splunk, Prometheus, Grafana, and strong troubleshooting across Node.js and Java services, with experience in CI/CD and GitOps tooling. This contract-based role is based in Montreal, Canada.

Qualifications

  • 5+ years in SRE/DevOps/platform engineering with strong Kubernetes ops.
  • 3+ years hands-on Kubernetes production experience, including deployment and troubleshooting.
  • Strong incident response experience with on-call, runbooks, and postmortems.
  • Splunk experience for log aggregation, search, and troubleshooting.
  • Prometheus and Grafana experience for building alerts and dashboards.
  • CI/CD and IaC for containerized deployments, including Helm and GitOps tools.

Responsibilities

  • Support deployment, operation, and reliability of production services on Kubernetes.
  • Monitor service health and investigate incidents across distributed apps.
  • Participate in on-call, incident response, RCA, and reliability improvements.
  • Troubleshoot runtime, networking, and inter-service issues with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within client-directed backlog and established priorities.

Job description

Company Description

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!

About the Client

Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.

You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.

Contract Duration

Initial contract through the end of 2026, extending the engagement to a total 12-month term based on performance.

Job Description
About the Role

This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.

This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership.

What You’ll Do
  • Support the deployment, operation, and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
  • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within a client-directed backlog and established priorities.
Qualifications
Required Qualifications
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
  • 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
  • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
  • Splunk experience for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
  • Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment. Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.
  • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication
Additional Information
Nice to Have
  • Web Components / Lit experience, to perform first-level debugging of UI-related issues
  • Server-side rendering or isomorphic runtime experience
  • Canary rollout / multi-version production operations
  • Distributed tracing and request-context correlation
  • KEDA or event-driven autoscaling
  • Experience with enterprise platform integration layers
What We Offer
  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 110,000 - 140,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Orion Innovation • British Columbia

On-site
CAD 100,000 - 130,000
Senior SRE: Kubernetes Reliability for Cloud UI Services
Senior SRE: Kubernetes Reliability for Cloud UI Services

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

On-site
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4
Senior SRE
Senior SRE

Viafoura • Toronto

Hybrid
CAD 100,000 - 130,000
Competitive salary
Comprehensive health benefits
Professional development opportunities
+2
DevOps, Kubernetes, and Site Reliability Engineer
DevOps, Kubernetes, and Site Reliability Engineer

Randstad Enterprise • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Senior Site Reliability Engineer, SRE
Senior Site Reliability Engineer, SRE

Jobtailor • Toronto

On-site
CAD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Platform & SRE Engineer
Platform & SRE Engineer

TechDoQuest • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
SRE x 2
SRE x 2

HRB • Montreal (administrative region)

On-site
CAD 110,000 - 170,000