Site Reliability Engineer

Future Secure AI

Toronto

On-site

CAD 90,000 - 130,000

Full time

43 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Future Secure AI is seeking a Site Reliability Engineer to design, build, and operate production platforms powering AI Co‑Workers. You will own end‑to‑end reliability, collaborating with product, AI, and engineering teams.

The role requires 5+ years in SRE/DevOps, Kubernetes expertise (EKS/AKS/GKE or self‑managed), and strong infrastructure as code experience with Terraform and Helm. Proficiency in Python/Go/Java/Bash/PowerShell/Ruby is expected, with CI/CD ownership and incident response

Qualifications

  • Bachelor's degree in computer science, information systems, or related field.
  • 5+ years of professional experience in Site Reliability/DevOps.
  • Kubernetes experience with EKS/AKS/GKE or self-managed clusters.
  • Terraform for provisioning and automation.
  • Helm for Kubernetes deployments.
  • Proficiency in Python, Go, Java, Bash, PowerShell, or Ruby.
  • Direct experience with reliability engineering, on-call rotations, and incident response.
  • CI/CD ownership and security considerations.

Responsibilities

  • Design, build, and operate reliable production infrastructure for AI workloads.
  • Own Kubernetes-based platforms used to deploy and run workloads.
  • Build and maintain infrastructure as code using Terraform.
  • Implement and maintain Helm-based deployment workflows.
  • Define, measure, and improve system reliability with SLIs, SLOs, and SLAs.
  • Participate in on-call rotation, incident response, root-cause analysis, and post-mortems.
  • Reduce operational toil through automation and engineering improvements.
  • Build and improve observability across monitoring, logging, and alerting.
  • Collaborate with engineers to ensure systems are resilient, scalable, and secure.
  • Operate across build, deploy, and operate phases of the software lifecycle.

Skills

Kubernetes
Terraform
Helm
Python
Go
Java
Bash
PowerShell
Ruby
CI/CD
On-call experience

Education

Bachelor's degree in Computer Science or related field
Masters Degree (preferred)

Tools

ArgoCD

Job description

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us. We work at the frontier of AI, tackling big, real-world problems for global enterprises across multiple industries, armed with state-of-the-art technology and a culture that prizes courage, rigor, and relentless curiosity. Our BRAVER values aren't just words on a wall — they describe the kind of people we are and the standard we hold ourselves to every day. Our leadership team is entrepreneurial, experienced, and accessible, with an open-door policy that means you'll never be just a number here. We invest seriously in your growth because we know our success depends on yours. If you're ready to work alongside some of the brightest minds in the industry, push into uncharted territory, and do work that genuinely matters, Future Secure AI is the place for you.

About the Role

We are looking for a Site Reliability Engineer to help design, build, and operate the platforms that power AI Co‑Workers. This is a hands‑on role for an engineer who enjoys owning reliability end‑to‑end and working closely with product, AI, and engineering teams.

Responsibilities
  • Design, build, and operate reliable production infrastructure supporting AI Co‑Workers
  • Own Kubernetes‑based platforms used to deploy and run AI workloads
  • Build and maintain infrastructure as code using Terraform
  • Implement and maintain Helm‑based deployment workflows
  • Define, measure, and improve system reliability using SLIs, SLOs, and SLAs
  • Participate in on‑call rotation, incident response, root cause analysis, and post‑mortems
  • Reduce operational toil through automation and engineering improvements
  • Build and improve observability across monitoring, logging, and alerting
  • Partner closely with engineers to ensure systems are resilient, scalable, and secure
  • Operate across build, deploy, and operate phases of the software lifecycle
Minimum Qualifications
  • Bachelors Degree in Computer Science, Information Systems, or related field
  • 5+ years of professional experience in Site Reliability Engineering or DevOps Engineering
  • Kubernetes experience designing, building, or operating workloads on EKS, AKS, GKE, or self‑managed Kubernetes
  • Terraform experience for infrastructure provisioning and automation
  • Helm experience for Kubernetes application deployment
  • Professional experience using at least two programming or scripting languages such as Python, Go, Java, Bash, PowerShell, or Ruby
  • Direct experience with reliability engineering, on‑call rotations, incident response, post‑mortems, and toil reduction
  • DevOps or DevSecOps experience, including CI/CD ownership, infrastructure automation, and security considerations
Preferred Qualifications
  • Masters Degree in Computer Science, Information Systems, or related field
  • Experience working within a defined SDLC, including CI/CD, release processes, and end‑to‑end delivery from design to operations
  • Hands‑on experience with at least one major cloud provider such as AWS, Azure, or Google Cloud
  • Experience with ArgoCD or GitOps‑style deployment approaches
  • Relevant certifications such as CKA, CKAD, cloud certifications, DevOps, DevSecOps, or programming credentials
Why Join Us?
  • A high-performance culture
  • State‑of‑the‑art technology
  • Experience world‑class leadership
  • Scale of impact and purpose
  • A competitive salary and a huge growth trajectory
  • Work with the best in the industry
  • Flexible work environment
  • Diversity and creativity
Disclaimer:

We do not wish to be contacted by recruitment agencies. Our hiring process is managed in-house.

Future Secure AI Privacy Policy

At Future Secure AI, we are committed to protecting your privacy and adhering to the principles of the General Data Protection Regulation (GDPR) and the Australian Privacy Principles (APPs) under the Privacy Act 1988 (Cth). Our Privacy Policy outlines how we collect, use, share, and protect your personal data when you visit our website atwww.futuresecure.ai(the "Website") and use our services.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Solutions Architect
Solutions Architect

Future Secure AI • Toronto

On-site
CAD 90,000 - 150,000
Flexible work environment
Growth opportunities
Site Reliability Engineer — Kubernetes & Terraform
Site Reliability Engineer — Kubernetes & Terraform

Future Secure AI • Toronto

On-site
CAD 90,000 - 130,000
Engagement Lead Toronto, CAN
Engagement Lead Toronto, CAN

Future Secure AI Pty • Toronto

On-site
CAD 100,000 - 130,000
Competitive salary
Performance incentives
Equity participation
+2
Quality Assurance Engineer
Quality Assurance Engineer

Future Secure AI • Toronto

On-site
CAD 75,000 - 95,000
Competitive salary
Flexible work environment
State-of-the-art technology
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AlleyCorp • Canada

On-site
CAD 197,000 - 225,000
Stock options
Health benefits
Unlimited PTO
+2
Software Engineer
Software Engineer

FutureFit • Canada

On-site
USD 150,000 - 185,000
Senior Site Reliability Developer
Senior Site Reliability Developer

United States Digital Space LLC • Toronto

On-site
CAD 107,000 - 157,000
Salary transparency
In-person onboarding
Senior SRE
Senior SRE

CloudFactory Limited • Canada

Hybrid
CAD 120,000 - 160,000
Hybrid Working Model
Comprehensive medical cover
Group life insurance
+3
Site Reliability Engineer
Site Reliability Engineer

RXinsider LTD. • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Baseten • Montreal (administrative region)

On-site
CAD 232,000 - 465,000
Competitive compensation & equity
Health, dental, vision insurance
Flexible PTO including Winter Break
+4