Founding Engineer (Site Reliability Engineer)

Katalyze AI, Inc.

Toronto

On-site

CAD 120,000 - 180,000

Full time

3 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Katalyze AI, Inc. is seeking a Founding Site Reliability Engineer to ensure the platform is reliable, scalable, and secure as we grow with enterprise customers in Toronto.

You will build and maintain infrastructure and practices that keep systems running smoothly while enabling rapid velocity. You will define and maintain SLOs/SLIs, design CI/CD pipelines, manage cloud infrastructure with IaC, implement observability, and partner with engineering to bake reliability into development.

Qualifications

  • 4+ years of SRE, DevOps, or platform engineering experience.
  • Strong experience with Kubernetes, Docker, and container orchestration.
  • Proficiency with cloud platforms (AWS preferred) and infrastructure-as-code.
  • Experience with observability tools (Datadog, Grafana, Prometheus, or similar).
  • Understanding of security best practices and enterprise compliance requirements (SOC 2, HIPAA awareness).
  • Experience with Python or Go for automation scripting.
  • Startup experience preferred — you're comfortable building from scratch.

Responsibilities

  • Define and maintain SLOs, SLIs, and error budgets for critical platform services.
  • Build and operate CI/CD pipelines, monitoring, alerting, and incident response systems.
  • Design and manage cloud infrastructure (AWS/GCP/Azure) using IaC (Terraform, Pulumi).
  • Implement observability tooling (logging, tracing, metrics) across the platform.
  • Partner with engineering to embed reliability practices into the development lifecycle.
  • Lead incident response and post-mortems; drive systemic improvements.
  • Support security and compliance requirements for enterprise customer deployments.
  • Build automation to reduce toil and improve operational efficiency.

Skills

SRE / DevOps experience
Kubernetes
Docker
container orchestration
cloud platforms (AWS)
infrastructure-as-code
observability tooling
security & compliance awareness
Python or Go
startup experience

Tools

Kubernetes
Docker
Terraform
Pulumi
Datadog
Grafana
Prometheus

Job description

About Katalyze AI

Katalyze AI is a fast-growing AI-driven biotech platform company on a mission to make life-saving drugs accessible and affordable for everyone. Our AI Agents help pharmaceutical and biotech companies increase production efficiency, reduce costs, and minimize waste. We're a team of humble, fast-moving, and curious craftspeople working at the intersection of science and AI.

About the Role

We're looking for a Founding Site Reliability Engineer to ensure Katalyze AI's platform is reliable, scalable, and secure as we grow with enterprise customers. You'll build and maintain the infrastructure and practices that keep our systems running smoothly and help us move fast without breaking things.

What You'll Do

Define and maintain SLOs, SLIs, and error budgets for critical platform services

Build and operate CI/CD pipelines, monitoring, alerting, and incident response systems

Design and manage cloud infrastructure (AWS/GCP/Azure) using infrastructure-as-code (Terraform, Pulumi)

Implement observability tooling (logging, tracing, metrics) across the platform

Partner with engineering to embed reliability practices into the development lifecycle

Lead incident response and post-mortems; drive systemic improvements

Support security and compliance requirements for enterprise customer deployments

Build automation to reduce toil and improve operational efficiency

What We're Looking For

4+ years of SRE, DevOps, or platform engineering experience

Strong experience with Kubernetes, Docker, and container orchestration

Proficiency with cloud platforms (AWS preferred) and infrastructure-as-code

Experience with observability tools (Datadog, Grafana, Prometheus, or similar)

Understanding of security best practices and enterprise compliance requirements (SOC 2, HIPAA awareness)

Experience with Python or Go for automation scripting

Startup experience preferred — you're comfortable building from scratch

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer (Agentic Platform)
Founding Engineer (Agentic Platform)

Katalyze AI, Inc. • Toronto

On-site
CAD 130,000 - 200,000
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Staff Data Engineer
Staff Data Engineer

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Mantu • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Forward Deployed Engineer (Staff/ Founding)
Forward Deployed Engineer (Staff/ Founding)

Katalyze AI, Inc. • Toronto

On-site
CAD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions Pvt Ltd • Toronto

On-site
CAD 120,000 - 170,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind Americas • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
Site Reliability Engineer — Kubernetes & Terraform
Site Reliability Engineer — Kubernetes & Terraform

Future Secure AI • Toronto

On-site
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Open Systems Technologies • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training