Founding Engineer (Agentic Platform)

Katalyze AI, Inc.

Toronto

On-site

CAD 130,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Katalyze AI, Inc. is seeking a Platform / Infrastructure Engineer to build and scale a secure, multi-tenant cloud backbone for our AI-powered pharma platforms. You will ensure SOC2/HIPAA compliance while delivering reliable, production-grade infrastructure with IaC, CI/CD, and observability.

You will own observability, security controls, and cost optimization, collaborating with teams across engineering to support global pharmaceutical customers and maintain high availability.

Qualifications

  • Experience managing production infrastructure in AWS (ECS/Fargate, RDS, S3, CloudFront, VPC networking).
  • Deep IaC patterns for multi-environment deployments (Dev/Staging/Prod).
  • Experience with Docker/ECS and auto-scaling health checks.
  • On-call incident response experience in production outages.
  • Strong CI/CD/automation for monorepos (Nx a plus).
  • Security: least-privilege IAM, network isolation, secrets management.

Responsibilities

  • Establish production-grade observability with OpenTelemetry and tracing to detect incidents quickly.
  • Implement automated provisioning and CI/CD pipelines for multi-tenant SaaS environments using Terraform and GitHub Actions.
  • Own platform reliability aiming for 99.9% uptime and fast API response times.
  • Enforce SOC2/HIPAA controls, manage secrets, IAM policies, and audit logging.
  • Drive cost optimization and right-size AWS resources.
  • Produce runbooks, postmortems, and ADRs for operational excellence.

Skills

AWS Mastery
Terraform Expertise
Containerization
Incident Response
CI/CD & Automation
Security Mindset
Overlap with US East Coast

Tools

Terraform
Docker
GitHub Actions
Nx
Airflow/MWAA
Snowflake
Kafka
Redis
BullMQ
dbt

Job description

Katalyze AI is a fast-growing AI-driven biotech platform company on a mission to make life-saving drugs accessible and affordable for everyone. Our AI Agents help pharmaceutical and biotech companies increase production efficiency, reduce costs, and minimize waste. We're a team of humble, fast-moving, and curious craftspeople working at the intersection of science and AI.

About the Role

We are looking for a Platform / Infrastructure Engineer to build and scale a reliable, secure, and multi-tenant infrastructure for our AI-powered pharmaceutical platforms. You will be responsible for the "production-grade" backbone of our services, ensuring that our systems meet strict SOC2/HIPAA compliance standards while serving global pharmaceutical leaders. From implementing OpenTelemetry to automating zero-downtime deployments with Terraform, you will own the reliability and security of our entire cloud ecosystem.

Key Responsibilities

Observability & Monitoring: Establish a production-grade observability stack (OpenTelemetry, distributed tracing) to achieve a <5 min mean-time-to-detection for incidents.

Infrastructure as Code (IaC): Implement automated provisioning and CI/CD pipelines (Terraform, GitHub Actions) for multi-tenant SaaS environments.

Reliability Engineering: Own platform metrics, aiming for a 99.9% uptime SLA and <500ms p95 API latency.

Security & Compliance: Implement and maintain SOC2/HIPAA controls, managing secrets (AWS Secrets Manager), IAM policies, and audit logging.

Cost Optimization: Drive architectural improvements and right-sizing strategies to optimize AWS expenditure.

Documentation: Create postmortems, runbooks, and architectural decision records (ADRs) to ensure team autonomy and operational excellence.

Qualifications

Required:

AWS Mastery: 5+ years managing production infrastructure (ECS/Fargate, RDS, S3, CloudFront, VPC networking).

Terraform Expertise: Deep experience with IaC patterns for multi-environment deployments (Dev/Staging/Prod).

Containerization: Battle-tested experience managing Docker/ECS with a focus on auto-scaling and health checks.

Incident Response: Real-world experience in on-call rotations and resolving live production outages.

CI/CD & Automation: Strong experience implementing pipelines for monorepo applications (Nx experience is a plus).

Security Mindset: Practical knowledge of least-privilege IAM, network isolation, and secrets management.

Overlap: Ability to work with at least 4-6 hours of overlap with US East Coast (EST/EDT) business hours.

Nice to Have:

  • Snowflake administration (role management, query optimization).
  • Python scripting for infra-automation.
  • Experience with Kafka, Redis, or BullMQ queue infrastructure.
  • Familiarity with dbt pipeline orchestration (Airflow/MWAA).
  • Infrastructure: Terraform, Docker, GitHub Actions.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer (Staff/ Founding)
Forward Deployed Engineer (Staff/ Founding)

Katalyze AI, Inc. • Toronto

On-site
CAD 140,000 - 180,000
Staff Platform Engineer
Staff Platform Engineer

Robots and Pencils • Calgary

Hybrid
CAD 96,000 - 138,000
Principal Engineer
Principal Engineer

Jobtailor • Toronto

On-site
CAD 180,000 - 240,000
Staff Data Engineer
Staff Data Engineer

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

MarkiTech.AI • Toronto

On-site
CAD 140,000 - 180,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Palona AI • Toronto

On-site
CAD 90,000 - 130,000
Competitive Salary
Stock Option Plan
Medical benefits
+4
Site Reliability Engineer
Site Reliability Engineer

Future Secure AI • Toronto

On-site
CAD 90,000 - 120,000
Flexible work environment
Competitive salary
Diversity and creativity
AI Platform Engineer – Senior/Principal
AI Platform Engineer – Senior/Principal

Jobtailor • Toronto

On-site
CAD 150,000 - 230,000
Senior ML Scientist
Senior ML Scientist

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Sr Platform Engineer
Sr Platform Engineer

Air-tek • Toronto

On-site
CAD 80,000 - 110,000