Site Reliability Engineer — Kubernetes & Terraform

Future Secure AI

Toronto

On-site

CAD 90,000 - 130,000

Full time

42 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Future Secure AI is seeking a Site Reliability Engineer to design, build, and operate production platforms powering AI Co‑Workers. You will own end‑to‑end reliability, collaborating with product, AI, and engineering teams.

The role requires 5+ years in SRE/DevOps, Kubernetes expertise (EKS/AKS/GKE or self‑managed), and strong infrastructure as code experience with Terraform and Helm. Proficiency in Python/Go/Java/Bash/PowerShell/Ruby is expected, with CI/CD ownership and incident response

Qualifications

  • Bachelor's degree in computer science, information systems, or related field.
  • 5+ years of professional experience in Site Reliability/DevOps.
  • Kubernetes experience with EKS/AKS/GKE or self-managed clusters.
  • Terraform for provisioning and automation.
  • Helm for Kubernetes deployments.
  • Proficiency in Python, Go, Java, Bash, PowerShell, or Ruby.
  • Direct experience with reliability engineering, on-call rotations, and incident response.
  • CI/CD ownership and security considerations.

Responsibilities

  • Design, build, and operate reliable production infrastructure for AI workloads.
  • Own Kubernetes-based platforms used to deploy and run workloads.
  • Build and maintain infrastructure as code using Terraform.
  • Implement and maintain Helm-based deployment workflows.
  • Define, measure, and improve system reliability with SLIs, SLOs, and SLAs.
  • Participate in on-call rotation, incident response, root-cause analysis, and post-mortems.
  • Reduce operational toil through automation and engineering improvements.
  • Build and improve observability across monitoring, logging, and alerting.
  • Collaborate with engineers to ensure systems are resilient, scalable, and secure.
  • Operate across build, deploy, and operate phases of the software lifecycle.

Skills

Kubernetes
Terraform
Helm
Python
Go
Java
Bash
PowerShell
Ruby
CI/CD
On-call experience

Education

Bachelor's degree in Computer Science or related field
Masters Degree (preferred)

Tools

ArgoCD

Job description

Future Secure AI is seeking a Site Reliability Engineer to design, build, and operate production platforms powering AI Co‑Workers. You will own end‑to‑end reliability, collaborating with product, AI, and engineering teams.

The role requires 5+ years in SRE/DevOps, Kubernetes expertise (EKS/AKS/GKE or self‑managed), and strong infrastructure as code experience with Terraform and Helm. Proficiency in Python/Go/Java/Bash/PowerShell/Ruby is expected, with CI/CD ownership and incident response

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Future Secure AI • Toronto

On-site
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Mantu • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Founding Engineer (Site Reliability Engineer)
Founding Engineer (Site Reliability Engineer)

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind Americas • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Senior SRE - Kubernetes Reliability for AI UI Platform
Senior SRE - Kubernetes Reliability for AI UI Platform

Software Mind Americas • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

TekRek • Vancouver

On-site
CAD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Open Systems Technologies • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Senior Site Reliability Engineer - Global Infra & CI/CD Impact
Senior Site Reliability Engineer - Global Infra & CI/CD Impact

CloudFactory Limited • Canada

Hybrid
CAD 120,000 - 160,000
Hybrid Working Model
Comprehensive medical cover
Group life insurance
+3