This role sits inside Cloud Engineering, the team responsible for the reliability, scalability, and operational maturity of PowerPlan's multi-cloud platform. We run production workloads across AWS and Azure for enterprise customers who don't tolerate downtime — which means our job is equal parts firefighting, systems thinking, and building the automation that means we stop having to firefight the same fire twice.
We're looking for a Principal Site Reliability Engineer to work hands-on across our AWS and Azure environments, solve the production problems that actually show up, and then engineer the toil out of them so they don't come back. This is a principal-level individual contributor role with real autonomy — you won't be told how to fix things, you'll be the one other engineers come to when they need to know how. Over the next year, you'll have the room to shape how reliability engineering is practiced across the whole organization, not just inside your own queue.
Success will be measured by how much manual, repetitive operational work you eliminate, how mature and calm our incident response becomes, and how much better on-call and engineering teams can see what's actually happening in production because of the observability platform you build.
- Resolve escalated infrastructure cases across major AWS and Azure services, and ship targeted Python or PowerShell automations against the patterns you find repeating.
- Analyze case and incident data to find the highest-frequency sources of operational toil, and eliminate or significantly reduce them through automation, self-service tooling, or infrastructure improvements.
- Lead critical incidents end-to-end, standardize incident runbooks, and facilitate blameless postmortems that actually produce tracked, completed corrective actions — not just a document nobody opens again.
- Design and build a mature observability layer across AWS and Azure — Grafana dashboards tied to real service health and user journeys, tuned alerts, and SLI/SLO reporting people actually use.
- Integrate metrics, logs, and traces from our core platforms, and embed observability directly into CI/CD and incident response workflows rather than bolting it on after the fact.
- Partner with Support, Professional Services, and Product to validate that the automations and tooling you build actually get adopted, not just deployed.
- Coach other engineers on effective incident communication and decision-making, and influence how reliability is practiced across teams you don't formally manage.
- You have deep, hands-on experience operating production systems in AWS and Azure environments — not just deploying to them, but keeping them alive under real pressure.
- You're strong with Python and/or PowerShell for operational automation, and you'd rather write a script once than fix the same ticket five times.
- You have a proven track record of spotting repetitive operational work and killing it with automation, self-service tooling, or infrastructure fixes.
- You've led incident response before — not just participated in it — and you know how to run a blameless postmortem that actually changes something.
- You have real observability chops, ideally with Grafana and SLI/SLO-driven monitoring, and you think in terms of service health and user journeys, not just dashboards for their own sake.
- You can influence how other engineers work without having formal authority over them — people take your recommendations seriously because your track record backs them up.
- You communicate clearly in writing and out loud, to engineers and non-technical stakeholders alike.
Preferred
- You've worked with infrastructure-as-code tooling (Terraform, Pulumi, or similar) to make reliability fixes repeatable, not one-off.
- You have exposure to Kubernetes or container orchestration in a production, multi-cloud context.
- You've operated in a compliance-driven environment (SOC 2 or similar) and know how that changes what "automation" is allowed to touch.
- You've used AI-assisted tooling or LLM-based approaches to speed up root-cause analysis, log parsing, or incident triage.
- You've built or contributed to an internal developer platform, self-service tooling, or a strong on-call/paging setup (PagerDuty, Opsgenie, or similar).
PowerPlan is an EOE
Applicant and Candidate Privacy Notice
Please note that this is a hybrid role that involves a combination of onsite work from our corporate office as well as work from home. While we strive to accommodate flexible working arrangements when sensible, there will be times when onsite work is required. This could include scheduled office days, team meetings, client meetings, or special events.