Site Reliability Engineer

rentsync_careers

Canada

Remote

CAD 130,000 - 170,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Rentsync is seeking a hands-on Site Reliability Engineer to lead incident response across our AWS and Kubernetes-based services in Canada. You’ll detect root causes, fix issues, and harden infrastructure while collaborating with multiple engineering teams.

You’ll own monitoring, automation, and on-call readiness, building reliable systems with Prometheus/Grafana, Terraform, and CI/CD pipelines, supporting 10+ products across cloud providers.

Qualifications

  • 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications.
  • Hands-on incident response experience where you diagnosed and fixed production issues yourself, not only coordinated them.
  • Strong, hands-on AWS experience in production (EKS, EC2, RDS, VPC networking, IAM, CloudWatch).
  • Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring Kubernetes workloads.
  • A track record of working with engineering teams to identify and resolve performance and reliability issues.
  • Experience with monitoring and observability tools (e.g. Prometheus, Grafana, Loki, Datadog, CloudWatch) and on-call/alerting tools such as PagerDuty.
  • Experience building automated tests or checks for production reliability (synthetics, smoke, health, or load testing).
  • Willingness to take part in an after-hours on-call rotation as we introduce one in the future.

Responsibilities

  • Be the first responder for production alerts and incidents across our services, and take them from triage through to resolution.
  • Diagnose and fix issues directly in AWS (EKS, EC2, RDS, networking, IAM) and Kubernetes, such as failing pods, resource exhaustion, bad deploys, networking/DNS, database and cache problems.
  • Roll back, scale, reconfigure, or patch infrastructure to restore service fast; elevate to development teams only when a code change is truly needed, and give them a clear diagnosis when you do.
  • Own our PagerDuty setup and incident response during business hours, driving down MTTD and MTTR.
  • Run blameless post-mortems and personally drive the technical follow-up work, not just the action-item list.
  • Automate runbooks and repetitive operational work, including using AI tools to speed up triage, investigation, and remediation.

Skills

Terraform
Linux
Networking
Scripting (Bash/Python)
Incident communication
CI/CD pipelines

Tools

Kubernetes
AWS (EKS, EC2, RDS)
CloudWatch
PagerDuty
Prometheus/Mimir
Grafana
Loki
OpenTelemetry
Tempo
GitHub Actions
GitLab CI

Job description

About Rentsync

Rentsync is an award-winning, high-growth organization that provides high quality websites, marketing services, and software solutions to the rental and property management industry throughout Canada and the United States.

About the role

We're looking for a hands-on Site Reliability Engineer to lead our response to production incidents. You'll dig into our AWS and Kubernetes environments, find the root cause, and fix it.

Between incidents, you'll make sure the same problem doesn't happen twice by hardening infrastructure, improving monitoring, and working with engineering teams on performance and reliability.

Our environment is bigger and more varied than most companies our size: 10+ products and 100+ services across multiple Kubernetes clusters, mainly on AWS with some Azure and GCP, built in PHP, Ruby on Rails, JavaScript/TypeScript, .NET, Python, and Rust.

Technologies you’ll work with:

AWS (EKS, EC2, RDS, S3, ALB/NLB, CloudWatch), Azure, GCP, Kubernetes, Terraform, Ansible, GitHub Actions/GitLab CI, PagerDuty, Prometheus/Mimir, Loki, Tempo, Grafana, OpenTelemetry, Cloudflare, Ubuntu & Amazon Linux, MySQL & PostgreSQL, Redis & Memcached, NGINX & Traefik, Bash, Python, and applications built in PHP, Ruby on Rails, JavaScript/TypeScript, .NET, and Rust.

This is a remote position.

Duties & Responsibilities
Incident response & first-contact remediation (primary focus)
  • Be the first responder for production alerts and incidents across our services, and take them from triage through to resolution.
  • Diagnose and fix issues directly in AWS (EKS, EC2, RDS, networking, IAM) and Kubernetes, such as failing pods, resource exhaustion, bad deploys, networking/DNS, database and cache problems.
  • Roll back, scale, reconfigure, or patch infrastructure to restore service fast; elevate to development teams only when a code change is truly needed, and give them a clear diagnosis when you do.
  • Own our PagerDuty setup and incident response during business hours, driving down MTTD and MTTR.
  • Run blameless post-mortems and personally drive the technical follow-up work, not just the action-item list.
  • Automate runbooks and repetitive operational work, including using AI tools to speed up triage, investigation, and remediation.
Reliability engineering (preventing the next incident)
  • Build and maintain monitoring for Kubernetes workloads and services (Prometheus/Mimir, Loki, Tempo, Grafana, OpenTelemetry), with low-noise, high-signal alerts.
  • Watch new releases in production and catch regressions in latency, errors, or resource use before they become incidents.
  • Create and maintain production test suites: synthetic checks, smoke tests, health checks, and load/performance tests.
  • Partner with engineering teams to find and fix performance and reliability issues, and define SLOs, SLIs, and error budgets.
  • Harden our platform after incidents: Terraform changes, Kubernetes resource tuning, autoscaling, CI/CD checks, secrets and IAM improvements.
  • Keep service docs and architecture decisions current so any engineer can operate our systems.
Required Knowledge, Skills & Abilities
  • Infrastructure as code with Terraform, and comfort working in CI/CD pipelines.
  • Solid Linux, networking, and container fundamentals.
  • Scripting/automation in Bash, Python, or similar.
  • Calm, clear communication during incidents and across teams
Essential Qualifications
  • 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications.
  • Hands-on incident response experience where you diagnosed and fixed production issues yourself, not only coordinated them.
  • Strong, hands-on AWS experience in production (EKS, EC2, RDS, VPC networking, IAM, CloudWatch).
  • Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring Kubernetes workloads.
  • A track record of working with engineering teams to identify and resolve performance and reliability issues.
  • Experience with monitoring and observability tools (e.g. Prometheus, Grafana, Loki, Datadog, CloudWatch) and on-call/alerting tools such as PagerDuty.
  • Experience building automated tests or checks for production reliability (synthetics, smoke, health, or load testing).
  • Willingness to take part in an after-hours on-call rotation as we introduce one in the future.
Additional Preferred Qualifications
  • Using AI tools to accelerate SRE work, such as incident triage, log and metric analysis, runbook automation, or infrastructure code.
  • Azure experience (GCP is a plus too).
  • Supporting many tech stacks across multiple teams (PHP, Ruby on Rails, .NET, Python, Rust, JavaScript).
  • AWS certification (e.g. Solutions Architect, DevOps Engineer, or SysOps).
  • Running LGTM (Loki, Grafana, Tempo, Mimir) or OpenTelemetry at scale.
  • Load testing tools such as k6, Locust, or JMeter.
  • MySQL/PostgreSQL operations and Redis/Memcached tuning.
  • Cloudflare (WAF, Workers, Zero Trust).
  • Cloud cost optimization and capacity planning.

Rentsync is an equal opportunity employer. If you are selected to participate in the interview process and require unique accommodations, please don't hesitate to let us know.

Successful candidates may be required to complete a criminal background check in the final phase of the interview process.

This is an open-ended job posting and may not represent a specific vacancy within the organization.

Rentsync reserves the right to use Artificial Intelligence to screen and/or assess candidates.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Rippling, Inc. • Toronto

On-site
CAD 110,000 - 165,000
Technical Support Specialist
Technical Support Specialist

rentsync_careers • St. Catharines

Hybrid
CAD 55,000 - 75,000
Senior .NET Developer
Senior .NET Developer

rentsync_careers • Montreal (administrative region)

Hybrid
CAD 110,000 - 140,000
Technical Support Specialist
Technical Support Specialist

Silversmith Capital Partners • Toronto

On-site
CAD 54,000 - 72,000
Senior .NET Developer
Senior .NET Developer

Rentsync • Montreal (administrative region)

Hybrid
CAD 100,000 - 120,000
Senior .NET Developer
Senior .NET Developer

Rentsync • Toronto

On-site
CAD 100,000 - 120,000
Hybrid work model
Technical Support Specialist
Technical Support Specialist

Rentsync • St. Catharines

Hybrid
CAD 50,000 - 58,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Rootly • Toronto

On-site
CAD 120,000 - 180,000
Competitive compensation
Comprehensive medical coverage
3 weeks of vacation
+3
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Relayfi • Toronto

Hybrid
CAD 153,000 - 187,000