DevOps, Kubernetes, and Site Reliability Engineer

Randstad Enterprise

Montreal (administrative region)

On-site

CAD 90,000 - 130,000

Full time

33 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Randstad Enterprise is seeking an experienced DevOps, Kubernetes, and Site Reliability Engineer to design, build, and operate scalable cloud-native platforms in Montreal. The role emphasizes automated deployment, strong scripting, and hands-on production support in a regulated environment.

The ideal candidate will have 5+ years in DevOps/SRE, solid Kubernetes and container experience, and a disciplined approach to reliability, security, and continuous improvement.

Qualifications

  • Minimum of 5 years of experience in DevOps, SRE, or production engineering.
  • Hands-on Kubernetes deployment, configuration, troubleshooting, and scaling.
  • Experience with Docker/Podman and containerized environments.
  • Strong Linux/UNIX administration and scripting (Python, Bash).
  • CI/CD concepts, release automation, and deployment strategies.
  • Experience with GitHub repositories, PRs, branching, and automated checks.
  • Proficiency with YAML, Ansible, or similar automation.
  • Solid understanding of networking (DNS, TCP/IP, HTTP, TLS, load balancing).
  • Experience in Agile environments and Jira familiarity.

Responsibilities

  • Design and maintain CI/CD pipelines and infrastructure.
  • Develop automated deployment solutions for Kubernetes environments.
  • Improve build, test, deployment, and release processes with development teams.
  • Support deployment, configuration changes, patches, and platform upgrades across environments.
  • Create and maintain Kubernetes deployment artifacts (YAML, Helm charts).
  • Automate tasks using Python and shell scripting; enhance automation tooling.
  • Manage GitHub repos, PR workflows, and access controls.
  • Integrate testing, security scanning, and linting into CI/CD pipelines.
  • Troubleshoot production issues across app, infra, container, and network layers.
  • Apply SRE principles to improve reliability and observability.
  • Define monitoring, alerting, dashboards, and runbooks.

Skills

Kubernetes
Docker/Podman
Linux
Networking
CI/CD
GitHub
Python
Shell scripting
SRE practices
Production support

Education

Bachelor's degree in CS/Engineering/IT or equivalent

Tools

Jenkins
Artifactory
Helm

Job description

We are looking for an experienced DevOps, Kubernetes, and Site Reliability Engineer with a minimum of five years of relevant industry experience. Experience within financial services or another highly regulated technology environment is preferred. The position requires strong hands-on experience with Kubernetes, Docker/Podman, Linux, networking, CI/CD pipelines, GitHub, Python, and shell scripting. The candidate should understand modern DevOps and SRE practices and be comfortable supporting production systems, troubleshooting complex application and infrastructure issues, and automating repetitive operational tasks.The ideal candidate will be passionate about automation, production reliability, and continuous improvement. The candidate should be organized, disciplined, detail-oriented, self-motivated, collaborative, and focused on delivering measurable engineering outcomes.

Location: Montreal (day 1 onboarding / onsite presence required 3x/week)

QUALIFICATIONS
Responsibilities
  • Design, build, maintain, and enhance CI/CD pipelines and supporting build, test, release, and deployment infrastructure.
  • Develop and maintain automated deployment solutions for applications running in Kubernetes and containerized environments.
  • Partner with application development teams to improve build, test, deployment, and release processes.
  • Support the deployment of applications, configuration changes, patches, and platform upgrades across development, testing, and production environments.
  • Build and maintain Kubernetes deployment artifacts, including YAML configuration, Helm charts, or equivalent packaging and configuration mechanisms.
  • Develop automation using Python, Linux shell scripting, and related tools.
  • Manage and improve GitHub repositories, branching strategies, pull-request workflows, access controls, and automated repository processes.
  • Integrate automated testing, code-quality validation, security scanning, dependency checks, and linting into CI/CD pipelines.
  • Troubleshoot application, infrastructure, container, Kubernetes, network, and deployment issues in complex environments.
  • Investigate production incidents, identify root causes, and implement preventive or corrective engineering solutions.
  • Apply SRE principles to improve system reliability, availability, scalability, observability, and operational readiness.
  • Define and improve monitoring, alerting, dashboards, operational metrics, and production support procedures.
  • Automate routine operational activities to reduce manual effort and operational risk.
  • Collaborate with infrastructure, network, database, cybersecurity, release management, and application teams.
  • Create and maintain technical documentation, operational runbooks, deployment procedures, and troubleshooting guides.
  • Participate in design reviews, production-readiness reviews, incident reviews, and continuous-improvement initiatives.
  • Participate in an after-hours production support and on-call rotation when required.
Required Skills
  • Minimum of 5 years of relevant experience in DevOps, SRE, production engineering, platform engineering, infrastructure engineering, or a related discipline.
  • Strong hands-on experience with Kubernetes, including application deployment, configuration, troubleshooting, scaling, services, ingress, secrets, and operational support.
  • Strong hands-on experience with Docker/Podman and containerized application environments.
  • Strong Linux and UNIX system administration and troubleshooting skills.
  • Strong experience with Linux shell scripting, such as Bash or KornShell.
  • Hands-on programming and automation experience using Python or a comparable language.
  • Strong understanding of CI/CD concepts, software delivery lifecycles, release automation, and deployment strategies.
  • Hands-on experience with CI/CD and artifact-management tools such as Jenkins and Artifactory, or equivalent platforms.
  • Hands-on experience with Git and GitHub, including repository management, pull requests, branching strategies, release workflows, and automated checks.
  • Experience developing automation using YAML, Ansible, or an equivalent automation framework.
  • Strong understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, TLS, proxies, firewalls, routing, load balancing, and network troubleshooting.
  • Experience supporting applications and resolving production issues in a fast-paced environment.
  • Understanding of SRE practices, including monitoring, incident response, root-cause analysis, service reliability, operational readiness, and continuous improvement.
  • Experience integrating code-quality tools, security scanning, automated testing, and policy controls into CI/CD pipelines.
  • Ability to troubleshoot issues across application, operating system, container, network, infrastructure, and database layers.
  • Experience working in an Agile development environment and using tools such as Jira.
  • Strong written and verbal communication skills.
  • Ability to collaborate effectively with globally distributed engineering and support teams.
  • Ability to prioritize work, manage multiple tasks, and deliver results with limited supervision.
  • Ability to participate in an after-hours on-call support rotation.
Desired Skills
  • Experience with enterprise Kubernetes platforms such as OpenShift or another managed Kubernetes environment.
  • Experience provisioning on-demand environments using virtual machines and containers.
  • Experience working with Azure, AWS, GCP, or another cloud platform.
  • Experience with infrastructure-as-code and configuration-management technologies.
  • Experience with Kubernetes package-management and deployment tools such as Helm.
  • Experience with GitOps deployment models and tools.
  • Knowledge of service mesh technologies, container networking, ingress controllers, and API gateways.
  • Experience with observability and telemetry platforms, including metrics, logs, traces, dashboards, and alerting.
  • Experience developing or maintaining Grafana dashboards.
  • Familiarity with production incident management, problem management, change management, service operations, and release management.
  • Understanding of high availability, disaster recovery, capacity management, and production resiliency.
  • Experience with relational database technologies such as DB2, Sybase, or Oracle.
  • Experience working in financial services or another regulated enterprise environment.
  • Familiarity with secure software supply-chain practices, secrets management, certificate management, and vulnerability remediation.
  • Education: Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical industry experience.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps, Kubernetes, and Site Reliability Engineer
DevOps, Kubernetes, and Site Reliability Engineer

Open Systems Technologies • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
DevOps / Site Reliability Engineer – Linux
DevOps / Site Reliability Engineer – Linux

Akkodis • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 110,000 - 140,000
DevOps & Site reliability Engineer
DevOps & Site reliability Engineer

TMC Canada • Montreal (administrative region)

On-site
CAD 110,000 - 140,000
Aucun
DevOps Architect (Remote - US or Canada)
DevOps Architect (Remote - US or Canada)

Jobgether • Canada

Remote
CAD 100,000 - 130,000
Competitive salary
Flexible work arrangements
Comprehensive healthcare coverage
+2
Java/Python Developer
Java/Python Developer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 70,000 - 90,000
Senior DevOps [#4977]
Senior DevOps [#4977]

Alteo Inc. • Montreal (administrative region)

Hybrid
CAD 110,000 - 150,000
Senior Observability Engineer
Senior Observability Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Montreal (administrative region)

Hybrid
CAD 120,000 - 160,000
Platform & SRE Engineer
Platform & SRE Engineer

TechDoQuest • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training