Staff Platform Engineer

Sage Recruiting Inc.

Canada

On-site

CAD 140,000 - 190,000

Full time

18 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Sage Recruiting Inc. is seeking a staff-level Platform Engineer to own end-to-end backend and cloud infrastructure for an AI-native startup. You will write production Go, design AWS/Kubernetes infra, and build Terraform/Helm tooling for hosted and customer deployments.

This senior IC role emphasizes reliability, enterprise readiness, and leadership in architectural decisions, with hands-on cloud and stack ownership in a fast-moving team.

Qualifications

  • You must have strong Go and backend development fundamentals, with end-to-end ownership.
  • Extensive cloud experience (AWS preferred) and deep Kubernetes expertise.
  • Proficient in Terraform, Helm, and CI/CD pipelines for reliable deployments.
  • Experience with distributed systems, observability, and database usage.
  • Proven ability to own ambiguous, critical problems end-to-end in production.

Responsibilities

  • Write and maintain production Go code for backend and services, from design to deployment.
  • Design and evolve AWS and Kubernetes infrastructure and deployment tooling.
  • Build reusable Terraform modules and Helm charts for hosted and customer deployments.
  • Harden product surfaces for enterprise readiness: auth, backups, availability.
  • Improve reliability with metrics, alerts, capacity planning, and on-call rotation.
  • Debug across the full stack and optimize CI/CD and development environments.
  • Lead technical design reviews, mentor teammates, and document decisions.

Skills

Go language
Cloud architecture
Kubernetes
Terraform
CI/CD
On-call
Distributed systems
AI-assisted workflows
Leadership/mentoring

Tools

Helm
Docker
PostgreSQL
OpenTelemetry
Prometheus

Job description

Sage Recruiting is partnering with an AI-native, early-stage infrastructure startup that's tackling a problem every engineering leader is quietly panicking about: AI coding agents now write code faster than any human team can review it, and the old ways of enforcing standards (wikis, checklists, manual review) were never built for that volume. Our client has built a guardrails engine that turns a company's engineering standards into automated enforcement, on every commit, every pull request, and every deploy, for human and AI-written code alike.

They're small, senior, and heavily AI-leveraged, moving with the speed and focus of a team several times their size. They've already landed their first paying customers and have real momentum building. Top-tier venture investors and an angel bench of well-known operators and creators from the developer tools world back them. This is a company at an inflection point between "early traction" and "real scale," actively going live with large enterprise customers, and this hire is one of the people who will build the infrastructure that the next chapter runs on.

The Role

This is a staff-level, deeply technical individual contributor role for a true backend/cloud/ops generalist. You'll be handed meaty, ambiguous problems and trusted to own them end-to-end: design, build, ship, and fix, with a seat in the on-call rotation like everyone else on the team. You'll write and own production Go code for the core platform, and you'll design and run the AWS and Kubernetes infrastructure it lives on, across both a hosted service and customer-managed deployments. A big part of the job is taking early, minimal product surfaces and making them enterprise-ready: authentication, backup, availability, and the reliability bar that large customers expect. You'll stay hands-on while helping decide what gets built and how, and you'll work directly with the company's forward-deployed engineers and customer platform teams on hard deployment and scale problems, feeding what you learn back into the product.

Who You Are:
  • You're a senior Platform Engineer who's comfortable working in code and cloud infrastructure.
  • You're comfortable being handed an ambiguous, high-stakes problem and can be trusted to run with it end-to-end without a fleshed-out playbook.
  • You think like a startup engineer: you know where the smart trade-offs are, where it's fine to cut a corner, and where it isn't, rather than defaulting to the most complete or "correct" solution.
  • You lean Kubernetes-strong in particular.
  • You're AI-native in your daily workflow, and you're excited by the idea of taking a product that's complete but still early and making it ready for large enterprise customers.
What You'll Do
  • Write and maintain production Go code in the backend and supporting services, owning features from design and testing through deployment and ongoing operation
  • Design and evolve the AWS architecture and Kubernetes infrastructure, making deliberate trade-offs around reliability, security, performance, cost, and how much complexity a small team can support
  • Build reusable Terraform modules, Helm charts, and deployment tooling that make provisioning, upgrades, and recovery repeatable, for the hosted service and customer-managed installations alike
  • Harden early-stage product surfaces to meet enterprise expectations: authentication, backups, availability, and the operational maturity large customers require
  • Improve production reliability: useful metrics and alerts, capacity planning, backups and tested recovery, safe rollouts, and clear rollback paths. Take your turn in the on-call rotation and fix the root causes of recurring problems, not just the symptoms
  • Debug across the full stack, following a failure through Go code, a database query, container behaviour, Kubernetes networking, or cloud infrastructure rather than stopping at a team boundary
  • Improve CI/CD and the development environment so engineers can test realistic changes and ship frequently without making production fragile
  • Partner with forward deployed engineers and customer platform teams on difficult deployment and scale problems, feeding what's learned back into the product
  • Lead technical design and code reviews, mentor teammates, and document decisions well enough that other engineers can maintain what you build
  • Build and manage AI-assisted engineering workflows for implementation, review, testing, and operational investigation, with appropriate access controls and checks on their output

You must have:

  • Strong Go and software engineering fundamentals: production services, concurrency, APIs, testing, performance work, and debugging distributed systems. This doesn't need to be your single deepest specialty, but you're comfortable owning backend code end-to-end
  • Deep cloud experience, ideally AWS (we're open to strong GCP or Azure backgrounds too): networking, IAM, compute, storage, and managed databases, with a real understanding of failure modes, isolation boundaries, and cost
  • Extensive, strong Kubernetes experience: built and operated production clusters and workloads, handled upgrades, debugged real failures, and worked with scheduling, networking, storage, RBAC, resource management, and Helm beyond an install guide. This is where we most want depth
  • Extensive Terraform experience: reusable modules, state, environment separation, drift, and safe changes to existing production infrastructure
  • Strong operational instincts: comfortable with Linux, Docker, networking, and troubleshooting under pressure, and you've owned systems after launch and made them easier to operate over time
  • Working knowledge of production data stores and observability: you can operate a SQL database (PostgreSQL or comparable) and object storage like S3, and you know how to instrument and read metrics, logs, and traces (Prometheus, Grafana, OpenTelemetry) well enough to actually run a system, not just build one
  • Extensive, hands-on use of AI coding tools and agents as part of your daily work. You build your own workflows, supply the context and tools they need, manage their permissions/cost/failure modes, and can explain how you verify what they produce
  • A track record of owning ambiguous, critical problems end-to-end and delivering changes that held up in production, whether that's 7+ years as a staff-level IC or a faster growth trajectory that's gotten you there
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer (Platform)
Senior Software Engineer (Platform)

Super • Toronto

Hybrid
CAD 140,000 - 190,000
Competitive base salary
Yearly learning & development Allow./s
RRSP & 401k employee contribution plan
+2
Software Engineering Team Lead
Software Engineering Team Lead

Exchange Solutions • Toronto

On-site
CAD 110,000 - 140,000
Founding Engineer (Agentic Platform)
Founding Engineer (Agentic Platform)

Katalyze AI, Inc. • Toronto

On-site
CAD 130,000 - 200,000
Senior Software Engineer, DevOps & Security
Senior Software Engineer, DevOps & Security

GuruLink • Vancouver

Hybrid
CAD 120,000 - 180,000
Hybrid work model
Downtown Vancouver office
Principal Engineer/ AI-Native Platform
Principal Engineer/ AI-Native Platform

Motion Recruitment • Toronto

On-site
CAD 170,000 - 240,000
Senior Software Engineer / Technical Lead
Senior Software Engineer / Technical Lead

Rise People • London

On-site
CAD 100,000 - 165,000
Staff Devops Engineer
Staff Devops Engineer

HRB • Ottawa

On-site
CAD 150,000 - 190,000
Director of Engineering (Infrastructure and Operations)
Director of Engineering (Infrastructure and Operations)

Motive • Vancouver

Hybrid
CAD 180,000 - 240,000
Senior Cloud Software Engineer
Senior Cloud Software Engineer

Harvey Nash • Vancouver

On-site
CAD 140,000 - 175,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

TekRek • Vancouver

On-site
CAD 150,000 - 210,000