Site Reliability Engineer

PagerDuty

Toronto

Hybrid

CAD 85,000 - 120,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Volunteer time off
Health insurance
Wellness days
Parental leave
Generous PTO
Flexible WFH
Career development programs

Job summary

PagerDuty is seeking a Site Reliability Engineer I on the Core Infrastructure team to help build and operate the foundational infrastructure powering PagerDuty’s real-time operations platform. You will work at the intersection of platform evolution and operational excellence, scaling and hardening networking, compute, and ingress across Kubernetes clusters while supporting global growth.

You’ll monitor health with metrics, logs, and on-call rotations, participate in agile rituals, and stay ahead

Qualifications

  • Networking fundamentals: load balancing, DNS, TLS, ingress traffic flow.
  • Linux-based production environments.
  • IaC with Terraform/CloudFormation; 0–1+ years in SRE/DevOps/PlatformEng roles.
  • Kubernetes/EKS experience.
  • Cloud-native infra across AWS/GCP/Azure.
  • Programming in Python, Ruby, or Go.
  • AWS networking concepts: VPCs, subnets, routing, security groups, load balancers.
  • Production Kubernetes platforms maintenance (EKS, ingress, networking).
  • Monitoring/observability tools: DataDog, New Relic, SumoLogic, Splunk, Prometheus, Grafana.
  • Service meshes, ingress controllers, API gateways: Envoy, Istio, NGINX.

Responsibilities

  • Build and operate foundational infrastructure powering PagerDuty’s platform.
  • Manage networking, compute, and ingress across clusters.
  • Scale and harden systems; contribute to reliability and security.

Skills

Networking fundamentals
Linux production systems
IaC (Terraform, CloudFormation)
Kubernetes / EKS
Cloud-native infra (AWS, GCP, Azure)
Programming (Python, Go, Ruby)
Load balancing / DNS / TLS / ingress
On-call / incident response
Monitoring & observability (Datadog, S

Tools

Terraform
CloudFormation
Kubernetes
AWS
GCP
Azure
Envoy / Istio / NGINX

Job description

  • As a Site Reliability Engineer I on the Core Infrastructure team you’ll help build and operate the foundational infrastructure that powers PagerDuty’s real-time digital operations platform
  • Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably
  • You’ll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems
  • Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases
  • Support and improve foundational infrastructure, including networking, compute
    platforms, Kubernetes clusters, and ingress/traffic management systems
  • Contribute to the reliability and scalability of PagerDuty’s core platform by
    hardening existing systems and supporting the rollout of new infrastructure
    capabilities
  • Participate in agile rituals (standups, planning, retros) and communicate
    progress/risks early
  • You stay current on technical trends to suggest innovative tools and approaches
    to interesting problems
  • Monitor system health using metrics, logs, and alerts, and participate in 24/7
    on-call rotations to help detect, respond to, and resolve incidents
Benefits
  • 20 hours per year of paid volunteer time
  • Health insurance
  • Wellness Days and mid-year Wellness Week: extra time off for whole company to unplug and recharge at the same time
  • Generous paid parental leave and return to work policy to help with transition back
  • Generous paid time off
  • Flexible workplace/WFH
  • Hands-on career and leadership development programs
Qualifications
  • Working knowledge of networking fundamentals, such as load balancing, DNS,TLS, and ingress traffic flow
  • Hands-on experience operating Linux-based systems in productionenvironments
  • Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)0 to 1+ years of experience in Site Reliability Engineering, DevOps, or PlatformEngineering roles
  • Experience with container orchestration (e.g., EKS, Kubernetes)
  • Experience working on cloud-native infrastructure (e.g., AWS, GCP, Azure),including networking and compute concepts
  • Proficiency in at least one programming language (e.g., Python, Ruby, Go, etc.)
  • Experience with AWS cloud networking concepts such as VPCs, subnets,routing, security groups, and load balancers
  • Experience operating or contributing to production Kubernetes platforms (e.g.,EKS), including cluster upgrades, networking, or ingress configuration
  • Experience with monitoring, observability, and logging platforms (e.g., DataDog,New Relic, SumoLogic, Splunk, Prometheus, Grafana)
  • Familiarity with service meshes, ingress controllers, or API gateways (e.g.,Envoy, Istio, NGINX)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer I
Site Reliability Engineer I

Pager • Toronto

Hybrid
CAD 136,000 - 206,000
Company equity
ESPP
Retirement plan
+4
Senior Site Reliability Engineer, SRE
Senior Site Reliability Engineer, SRE

Jobtailor • Toronto

On-site
CAD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 110,000 - 140,000
Staff, Site Reliability Engineer(Global Security)
Staff, Site Reliability Engineer(Global Security)

rbc • Toronto

On-site
CAD 140,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

twentysix • Vancouver

On-site
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000
Senior Infrastructure SRE
Senior Infrastructure SRE

PointClickCare • Mississauga

On-site
CAD 110,000 - 150,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Markham

On-site
CAD 120,000 - 180,000
Staff Developer - Incident Command
Staff Developer - Incident Command

IBM • Calgary

On-site
CAD 140,000 - 190,000
Senior Observability Engineer
Senior Observability Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Montreal (administrative region)

Hybrid
CAD 120,000 - 160,000