Senior Infrastructure SRE

PointClickCare

Mississauga

Hybrid

CAD 110,000 - 150,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

PointClickCare builds cloud platforms that power safer, more connected care for millions of patients. You'll design and operate resilient infrastructure services across Azure, AWS, GCP, and on‑prem environments, driving reliability, automation, and operational excellence for identity, compute, storage, messaging, and shared services with an SRE‑first mindset.

You will design scalable IaC, automate toil, and lead incident response, mentoring junior SREs and partnering with product engineering

Qualifications

  • 5+ years of hands-on experience operating and designing cloud infrastructure.
  • Expert-level understanding of Azure or AWS.
  • Multi-cloud experience with Azure, AWS and/or GCP.
  • 3+ years of production IaC experience (Terraform, Pulumi, CloudFormation).
  • Ability to design scalable IaC modules and enforce GitOps workflows.
  • Strong programming proficiency (Python, Go, Bash) for automation.
  • Demonstrated depth across Kubernetes and containerized workloads in production.
  • Experience with service mesh (Istio) in production.
  • Proven ability to define SLIs/SLOs, error budgets, and toil metrics.
  • Knowledge of identity/IAM and federation platforms.

Responsibilities

  • Design and implement highly available infrastructure solutions for compute, storage, identity, messaging, and shared services
  • Build and maintain Infrastructure as Code using Terraform or Pulumi; establish best practices and standards
  • Automate operational workflows to eliminate toil: auto-remediation, self-healing systems, capacity planning
  • Define and track SLIs and SLOs for critical services; manage error budgets and reliability targets
  • Participate in the on-call rotation and lead incident response for complex infrastructure issues; conduct blameless post-mortems and drive systemic fixes
  • Develop observability strategies: metrics, logs, distributed tracing, alerting frameworks
  • Apply AI-assisted tooling to reduce toil and speed up investigation: log analysis, alert triage, runbook and post-mortem drafting, automation scaffolding
  • Own reliability for one or more infrastructure domains end to end, driving multi-team initiatives with product engineering from problem definition through adoption
  • Mentor intermediate SREs; review infrastructure changes; establish operational best practices

Skills

Azure/AWS expertise
Multi-cloud proficiency
IaC (Terraform/Pulumi)
GitOps workflows
Python/Go/Bash
Automation scripting
SRE metrics (SLIs/SLOs)
Kubernetes production
Service mesh (Istio)
Identity/IAM
Federation platforms (PingFederate/En"

Tools

Terraform
Pulumi
CloudFormation
Kubernetes (AKS/EKS)
Istio
Git

Job description

At PointClickCare our mission is simple: to help providers deliver exceptional care. And that starts with our people. As a leading health tech company that’s founder-led and privately held, we empower our employees to push boundaries, innovate, and shape the future of healthcare.

With the largest long-term and post-acute care dataset and a Marketplace of 400+ integrated partners, our platform serves over 30,000 provider organizations, making a real difference in millions of lives. We also reinvest a significant percentage of our revenue back into research and development, ensuring our employees have the resources to innovate and make a lasting impact.Recognized by Forbes as a top private cloud company and honored as one of Canada’s Most Admired Corporate Cultures, we offer flexibility, growth opportunities, and meaningful work.

At PointClickCare, we empower our people to be the architects of a smarter healthcare future; one that is human-first and accelerated by AI to create meaningful and lasting change. Employees harness AI as a catalyst for creativity, productivity, and thoughtful decision-making. By integrating AI tools into our daily workflows, collaboration is enhanced, outcomes are improved, and every team member has the proficiency to maximize their impact. It all starts with our hiring practices where we uncover AI expertise that complements our mission, and we continue to invest in training and development to nurture innovation throughout the employee journey.

Join us in redefining healthcare — so it doesn’t just survive, it thrives.To learn more about PointClickCare, check out Life at PointClickCare and connect with us on Glassdoor and LinkedIn.

**Travel to Office expectations**
For Remote Roles

If this role is remote, there will be in-office events that will require travel to and from the Mississauga and/or Salt Lake City office. These will include, but not limited to, onboarding, team events, semi-annual and annual team meetings.

For Hybrid Roles

If this role is Hybrid, there will be an expectation to reside within commutable distance to the office/location specified in the job listing. This will include, but not limited to, weekly/bi-weekly/monthly events in the office with your specific team. This is a requirement for this role.

About the role

PointClickCare builds cloud platforms that power safer, more connected care for millions of patients. You'll design and operate resilient infrastructure services across Azure, AWS, GCP, and on-prem environments, driving reliability, automation, and operational excellence for identity, compute, storage, messaging, and shared services with an SRE-first mindset.

What you'll do
  • Design and implement highly available infrastructure solutions for compute, storage, identity, messaging, and shared services
  • Build and maintain Infrastructure as Code using Terraform or Pulumi; establish best practices and standards
  • Automate operational workflows to eliminate toil: auto-remediation, self-healing systems, capacity planning
  • Define and track SLIs and SLOs for critical services; manage error budgets and reliability targets
  • Participate in the on-call rotation and lead incident response for complex infrastructure issues; conduct blameless post-mortems and drive systemic fixes
  • Develop observability strategies: metrics, logs, distributed tracing, alerting frameworks
  • Apply AI-assisted tooling to reduce toil and speed up investigation: log analysis, alert triage, runbook and post-mortem drafting, automation scaffolding
  • Own reliability for one or more infrastructure domains end to end, driving multi-team initiatives with product engineering from problem definition through adoption
  • Mentor intermediate SREs; review infrastructure changes; establish operational best practices
Must-haves
  • 5+ years of hands-on experience operating and designing cloud infrastructure
  • Expert-level understanding of the core services of either Azure or AWS
  • Working proficiency in at least one additional platform (Azure, AWS, or GCP)
  • Experience designing and supporting production infrastructure that spans multiple cloud platforms
  • 3+ years of production experience with Infrastructure as Code (Terraform, Pulumi, CloudFormation)
  • Ability to design scalable, reusable IaC modules and enforce GitOps workflows
  • Experience managing IaC across multiple cloud providers, including module design, state layout, and provider-specific resource differences
  • Strong proficiency in at least one programming language (Python, Go, Bash) for production automation
  • Demonstrated ability to write tested, maintainable automation and tooling
  • Practical application of SRE principles in production environments
  • Experience defining and managing SLIs/SLOs, error budgets, toil metrics
  • Track record of improving system reliability (e.g., MTTR reduction, availability improvements)
  • Demonstrated depth across key infrastructure services:
  • Expert-level experience running Kubernetes and containerized workloads in production on managed Kubernetes (AKS, EKS, or equivalent), plus VM-based compute
  • Practical experience operating a service mesh in production (Istio preferred; Linkerd or equivalent)
  • Strong proficiency with enterprise identity and SSO: SAML, OAuth/OIDC, LDAP, and cloud IAM
  • Working knowledge of an enterprise federation platform such as PingFederate, Entra ID, Okta, or ADFS
  • Strong proficiency in storage solutions (object, block, and file storage; Kubernetes persistent volumes)
  • Working knowledge of designing and operating infrastructure in a regulated environment (HIPAA, SOC 2, PCI, FedRAMP, or equivalent)
  • Familiarity with audit evidence, access controls, encryption in transit and at rest, and data residency constraints
  • Proven track record of measurably reducing operational toil through automation (e.g. ticket volume, manual runbook executions, hours reclaimed)
  • Strong communication and documentation skills; demonstrated ability to influence engineering teams
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infrastructure SRE
Senior Infrastructure SRE

Jobtailor • Mississauga

On-site
CAD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Senior Site Reliability Engineer, SRE
Senior Site Reliability Engineer, SRE

Jobtailor • Toronto

On-site
CAD 120,000 - 180,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Saviynt • Vancouver

On-site
CAD 140,000 - 210,000
Senior Software Engineer, Infrastructure
Senior Software Engineer, Infrastructure

Jobtailor • Toronto

On-site
CAD 110,000 - 160,000
Senior Observability Engineer
Senior Observability Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Montreal (administrative region)

Hybrid
CAD 120,000 - 160,000
Platform Engineer
Platform Engineer

Swoon • Toronto

On-site
CAD 69,000 - 96,000
Site Reliability Engineer
Site Reliability Engineer

TELUS Digital • Canada

Remote
CAD 90,000 - 120,000
Senior Software Engineer, Platform Engineering
Senior Software Engineer, Platform Engineering

PointClickCare • Mississauga

On-site
CAD 90,000 - 120,000