Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc.

Toronto

On-site

CAD 120,000 - 180,000

Full time

3 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Katalyze AI, Inc. is seeking a Founding Site Reliability Engineer to ensure the platform is reliable, scalable, and secure as we grow with enterprise customers in Toronto.

You will build and maintain infrastructure and practices that keep systems running smoothly while enabling rapid velocity. You will define and maintain SLOs/SLIs, design CI/CD pipelines, manage cloud infrastructure with IaC, implement observability, and partner with engineering to bake reliability into development.

Qualifications

  • 4+ years of SRE, DevOps, or platform engineering experience.
  • Strong experience with Kubernetes, Docker, and container orchestration.
  • Proficiency with cloud platforms (AWS preferred) and infrastructure-as-code.
  • Experience with observability tools (Datadog, Grafana, Prometheus, or similar).
  • Understanding of security best practices and enterprise compliance requirements (SOC 2, HIPAA awareness).
  • Experience with Python or Go for automation scripting.
  • Startup experience preferred — you're comfortable building from scratch.

Responsibilities

  • Define and maintain SLOs, SLIs, and error budgets for critical platform services.
  • Build and operate CI/CD pipelines, monitoring, alerting, and incident response systems.
  • Design and manage cloud infrastructure (AWS/GCP/Azure) using IaC (Terraform, Pulumi).
  • Implement observability tooling (logging, tracing, metrics) across the platform.
  • Partner with engineering to embed reliability practices into the development lifecycle.
  • Lead incident response and post-mortems; drive systemic improvements.
  • Support security and compliance requirements for enterprise customer deployments.
  • Build automation to reduce toil and improve operational efficiency.

Skills

SRE / DevOps experience
Kubernetes
Docker
container orchestration
cloud platforms (AWS)
infrastructure-as-code
observability tooling
security & compliance awareness
Python or Go
startup experience

Tools

Kubernetes
Docker
Terraform
Pulumi
Datadog
Grafana
Prometheus

Job description

Katalyze AI, Inc. is seeking a Founding Site Reliability Engineer to ensure the platform is reliable, scalable, and secure as we grow with enterprise customers in Toronto.

You will build and maintain infrastructure and practices that keep systems running smoothly while enabling rapid velocity. You will define and maintain SLOs/SLIs, design CI/CD pipelines, manage cloud infrastructure with IaC, implement observability, and partner with engineering to bake reliability into development.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer (Site Reliability Engineer)
Founding Engineer (Site Reliability Engineer)

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Senior SRE Leader: Scale Reliability & Observability
Senior SRE Leader: Scale Reliability & Observability

Rootly • Toronto

On-site
CAD 120,000 - 180,000
Competitive compensation
Comprehensive medical coverage
3 weeks of vacation
+2
Senior Site Reliability Engineer - Global Infra & CI/CD Impact
Senior Site Reliability Engineer - Global Infra & CI/CD Impact

CloudFactory Limited • Canada

Hybrid
CAD 120,000 - 160,000
Hybrid Working Model
Comprehensive medical cover
Group life insurance
+3
Senior Cloud SRE & Reliability Leader
Senior Cloud SRE & Reliability Leader

Jobgether • Canada

Hybrid
CAD 151,000 - 200,000
Health benefits
Equity stock options
Remote-friendly environment
+2
Senior Site Reliability Engineer – Cloud & Automation Lead
Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc. • Toronto

Remote
CAD 90,000 - 120,000
Digital-first work environment
Collaborative workspaces
Continuous learning opportunities
Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
Senior SRE: Global SaaS Platform, Kubernetes & Cloud
Senior SRE: Global SaaS Platform, Kubernetes & Cloud

Kong Inc. • Toronto

On-site
CAD 100,000 - 130,000
Founding SRE Lead for Scalable Fintech Platform
Founding SRE Lead for Scalable Fintech Platform

Sage Recruiting Inc. • Canada

On-site
CAD 180,000 - 200,000
Senior SRE — Distributed Data Platforms & Cloud Infra
Senior SRE — Distributed Data Platforms & Cloud Infra

OpenText • Southwestern Ontario

On-site
CAD 80,000 - 131,000