Staff Site Reliability Engineer

Johnson Controls, Inc.

Richmond Hill

On-site

CAD 120,000 - 180,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Johnson Controls, Inc. is seeking a Staff Site Reliability Engineer in Canada to own escalations for the OpenBlue Data Platform and Airwall. You will lead end-to-end incident fixes, plan infrastructure upgrades, and ensure platform reliability across Azure and AWS.

You will work with Terraform as code, participate in on‑call rotations, and collaborate with security and compliance teams to protect sensitive government and commercial data.

Qualifications

  • 7+ years in site reliability engineering, production engineering, L3 application support, or infrastructure engineering
  • Demonstrated ability to take a complex production problem end to end and land a permanent fix
  • Strong hands on Terraform, with real production ownership of infrastructure as code
  • Production experience across Azure and AWS, and with Kubernetes at scale
  • Practical depth in Datadog and Grafana, covering instrumentation, dashboards, monitor design, and alert quality
  • Experience with AI native development, knowing when and when not to use harnesses
  • Willingness to participate in an on‑call rotation and to respond to production incidents outside business hours
  • Must reside in Canada and be legally authorized to work in Canada without sponsorship

Responsibilities

  • Serve as the senior escalation owner for production issues across OpenBlue Data Platform and Airwall
  • Debug complex cross-layer failures spanning app code, data pipelines, cloud infrastructure, and network paths
  • Lead root cause analysis and drive permanent corrective action with owning teams
  • Turn recurring escalations into engineering work by feeding defects back into the product backlog
  • Improve detection with instrumentation, monitors, dashboards, and runbooks in Datadog and Grafana
  • Participate in on‑call rotation and act as senior technical lead during major incidents
  • Represent Engineering with customers during high severity incidents and post‑incident reviews
  • Plan and execute infrastructure upgrades and migrations across Azure and AWS with minimal disruption

Skills

Terraform
Kubernetes
Azure
AWS
Datadog
Grafana
Java
C#
AI-native development
On-call

Tools

Terraform
Azure
AWS
Kubernetes
Datadog
Grafana

Job description

About Johnson Controls

Johnson Controls,a global leader in thermal management, mission-critical building systems, energy efficiency, and decarbonization, helps customers use energy more productively, reduce carbon emissions, and operate with the precision and resilience required in rapidly expanding industries such as datacenters, healthcare, pharmaceuticals, advanced manufacturing, and higher education.

For more than 140 years, Johnson Controls has delivered performance where itreally matters. Backed by advanced technology, lifecycle services and an industry-leading field organization, weelevatecustomer performance,turn goals into real-world results andhelpmove society forward.

Visit johnsoncontrols.com for more information and follow @Johnsoncontrols on socialplatforms.

What you will do

OpenBlue from Johnson Controls is a cyber-secured smart building ecosystem that unifies data, AI, and automation to transform how buildings perform. By connecting systems that have historically stood apart and applying award-winning analytics, we give customers real-time visibility, predictive insight, and automated action across the entire building lifecycle. None of that reaches a customer without the platform underneath it. Our data platform, our enterprise SaaS portfolio, and OpenBlue Airwall run continuously for enterprise, public sector, and government customers, in environments where an outage or a data integrity problem carries real operational consequence for the buildings and the people inside them.

Johnson Controls is seeking a Staff Site Reliability Engineer. You will be the senior technical owner of escalated production problems, capable of debugging a failure across application code, data pipelines, cloud services, and network paths, and driving it through to permanent corrective action rather than a restart and a hopeful note in the ticket. You will also plan and execute the infrastructure work that keeps those platforms healthy, with Terraform as your native language and change safety as your standing constraint. You will be a core team member of our engineering department, and participate in the on‑call rotation, and occasionally join customer conversations when the severity of an issue warrants an engineer in the room.

This role is based in Canada. We support Canadian government customers and commercial customers with Canadian data residency requirements, and this position works directly with those systems and their data.

How you will do it
Application reliability and L3 escalation
  • Serve as the senior escalation owner for production issues across theOpenBlue Data Platform, our enterprise SaaS products, and Airwall

  • Debug complex, cross layer failures spanning application code, data pipelines, cloud infrastructure, and network paths

  • Lead root cause analysis and drive both interim and permanent corrective action to closure with the owning engineering teams, including the code or configuration change that prevents recurrence

  • Turn recurring escalations into engineering work by feeding defect patterns, reliability gaps, and supportability problems back into the product backlog

  • Improve detection ahead of the customer by strengthening instrumentation, monitors, dashboards, and runbooks in Datadog and Grafana

  • Participate in the on‑call rotation and act as a senior technical lead during major incidents

  • Represent Engineering directly with customers during high severity incidents and post incident reviews when the situation calls for it

Infrastructure and platform engineering
  • Plan and execute infrastructure upgrades, migrations, and platform changes across Azure and AWS with minimal customer disruption

  • Own infrastructure as code in Terraform, including module design, state management, and drift remediation

  • Operate and improve Kubernetes workloads across capacity, autoscaling, resource limits, and deployment reliability

  • Raise release and change safety, and reduce manual toil through automation

  • Harden the environments serving government and data residency sensitive customers, partnering with security and compliance on controls and evidence

  • Set operational standards for observability, change management, and production readiness that other engineering teams adopt

What you will need
Required
  • Must reside in Canada and be legally authorized to work in Canada without sponsorship. This requirement supports our Canadian government customers and customers with Canadian data residency obligations

  • 7+ years in site reliability engineering, production engineering, L3 application support, or infrastructure engineering

  • Demonstrated ability to take a complex production problem end to end and land a permanent fix, with examples you can walk through

  • Strong hands on Terraform, with real production ownership of infrastructure as code

  • Production experience across Azure and AWS, and with Kubernetes at scale

  • Practical depth in Datadog and Grafana, covering instrumentation, dashboards, monitor design, and alert quality

  • Experience with AI native development, knowing when and when not to use harnesses such as Claude, Copilot, Codex, or Cursor as a core part of your workflow

  • Willingness to participate in an on‑call rotation and to respond to production incidents outside business hours

Preferred
  • Working proficiency in Java and C#, sufficient to read, diagnose, and correct application code

  • Experience supporting government or public sector customers, including data residency, data sovereignty, or Protected B handling requirements

  • Eligible to obtain Government of Canada security screening at Reliability Status or higher

  • Experience operating data platforms, including streaming and batch pipelines, data quality monitoring, and latency service level objectives

  • Familiarity with operational technology networking and zero trust network architecture

  • Prior experience in an L2 or L3 support organization with formal service level agreements and escalation structures

  • Exposure to building automation, HVAC, or connected building technology

#LI-ONSITE

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Tecsys Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 120,000
Digital-first work environment
Collaborative workspaces
Continuous learning opportunities
Senior Site Reliability Developer
Senior Site Reliability Developer

United States Digital Space LLC • Toronto

On-site
CAD 107,000 - 157,000
Salary transparency
In-person onboarding
Site Reliability Engineer / Full Stack Engineer
Site Reliability Engineer / Full Stack Engineer

Lockheed Martin • Dartmouth

On-site
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

RXinsider LTD. • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Staff Database Engineer
Staff Database Engineer

ServiceNow • Toronto

On-site
CAD 100,000 - 150,000
Senior DevOps Engineer
Senior DevOps Engineer

Quest Global • Vancouver

On-site
CAD 100,000 - 120,000
401(k) matching
Health insurance
Dental insurance
+5
Manager, Systems Engineering
Manager, Systems Engineering

ServiceNow • Toronto

On-site
CAD 126,000 - 220,000
Health plans
RRSP with company match
ESPP
+3
Snr Cloud Infrastructure Engineer
Snr Cloud Infrastructure Engineer

Bosa Properties Inc • Vancouver

On-site
CAD 105,000 - 143,000
SRE specialist
SRE specialist

Intact Financial Corporation • Montreal (administrative region)

On-site
CAD 109,000 - 135,000
Flexible work arrangements
Possibility to purchase up to 5 extra days off
Wellness benefits including telemedicine
+1
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000