Principal Site Reliability Engineer

PowerPlan, Inc.

Atlanta (GA)

Hybrid

USD 160,000 - 230,000

Full time

21 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

PowerPlan, Inc. is seeking a Principal Site Reliability Engineer to own and improve reliability across AWS and Azure in a hybrid onsite/remote setting in Atlanta.

You will automate toil away, design an observability platform, and lead critical incidents end-to-end while coaching peers and shaping practices across the org. You will apply deep production experience, strong scripting skills in Python/PowerShell, and a track record of reducing operational toil through self-service tooling and robust

Qualifications

  • Experience operating production systems in AWS and Azure.
  • Strong Python and/or PowerShell for automation.
  • Proven incident response experience with blameless postmortems.

Responsibilities

  • Resolve escalated infrastructure cases and ship automation to remove toil.
  • Lead incidents end-to-end and standardize runbooks with corrective actions.
  • Design a mature observability layer using Grafana with SLI/SLO reporting.
  • Integrate metrics, logs, and traces into CI/CD and incident workflows.
  • Collaborate with Support, Professional Services, and Product to drive adoption of tooling.
  • Mentor engineers on incident communication and decision-making.

Skills

Python scripting
PowerShell
Automation thinking
Incident management
Communication

Tools

Grafana
Terraform
Pulumi
Kubernetes
PagerDuty / Opsgenie

Job description

This role sits inside Cloud Engineering, the team responsible for the reliability, scalability, and operational maturity of PowerPlan's multi-cloud platform. We run production workloads across AWS and Azure for enterprise customers who don't tolerate downtime — which means our job is equal parts firefighting, systems thinking, and building the automation that means we stop having to firefight the same fire twice.

We're looking for a Principal Site Reliability Engineer to work hands-on across our AWS and Azure environments, solve the production problems that actually show up, and then engineer the toil out of them so they don't come back. This is a principal-level individual contributor role with real autonomy — you won't be told how to fix things, you'll be the one other engineers come to when they need to know how. Over the next year, you'll have the room to shape how reliability engineering is practiced across the whole organization, not just inside your own queue.

Success will be measured by how much manual, repetitive operational work you eliminate, how mature and calm our incident response becomes, and how much better on-call and engineering teams can see what's actually happening in production because of the observability platform you build.

  • Resolve escalated infrastructure cases across major AWS and Azure services, and ship targeted Python or PowerShell automations against the patterns you find repeating.
  • Analyze case and incident data to find the highest-frequency sources of operational toil, and eliminate or significantly reduce them through automation, self-service tooling, or infrastructure improvements.
  • Lead critical incidents end-to-end, standardize incident runbooks, and facilitate blameless postmortems that actually produce tracked, completed corrective actions — not just a document nobody opens again.
  • Design and build a mature observability layer across AWS and Azure — Grafana dashboards tied to real service health and user journeys, tuned alerts, and SLI/SLO reporting people actually use.
  • Integrate metrics, logs, and traces from our core platforms, and embed observability directly into CI/CD and incident response workflows rather than bolting it on after the fact.
  • Partner with Support, Professional Services, and Product to validate that the automations and tooling you build actually get adopted, not just deployed.
  • Coach other engineers on effective incident communication and decision-making, and influence how reliability is practiced across teams you don't formally manage.
  • You have deep, hands-on experience operating production systems in AWS and Azure environments — not just deploying to them, but keeping them alive under real pressure.
  • You're strong with Python and/or PowerShell for operational automation, and you'd rather write a script once than fix the same ticket five times.
  • You have a proven track record of spotting repetitive operational work and killing it with automation, self-service tooling, or infrastructure fixes.
  • You've led incident response before — not just participated in it — and you know how to run a blameless postmortem that actually changes something.
  • You have real observability chops, ideally with Grafana and SLI/SLO-driven monitoring, and you think in terms of service health and user journeys, not just dashboards for their own sake.
  • You can influence how other engineers work without having formal authority over them — people take your recommendations seriously because your track record backs them up.
  • You communicate clearly in writing and out loud, to engineers and non-technical stakeholders alike.
Preferred
  • You've worked with infrastructure-as-code tooling (Terraform, Pulumi, or similar) to make reliability fixes repeatable, not one-off.
  • You have exposure to Kubernetes or container orchestration in a production, multi-cloud context.
  • You've operated in a compliance-driven environment (SOC 2 or similar) and know how that changes what "automation" is allowed to touch.
  • You've used AI-assisted tooling or LLM-based approaches to speed up root-cause analysis, log parsing, or incident triage.
  • You've built or contributed to an internal developer platform, self-service tooling, or a strong on-call/paging setup (PagerDuty, Opsgenie, or similar).

PowerPlan is an EOE

Applicant and Candidate Privacy Notice

Please note that this is a hybrid role that involves a combination of onsite work from our corporate office as well as work from home. While we strive to accommodate flexible working arrangements when sensible, there will be times when onsite work is required. This could include scheduled office days, team meetings, client meetings, or special events.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Plano (TX)

On-site
USD 152,000 - 192,000
Discretionary incentive eligible
Benefits eligible
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Plano (TX)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Paid time off
Access to resources and support
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Charlotte (NC)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Paid time off
Discretionary incentive eligibility
Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Koitecc Solutions • Plano (TX), Northern (KY)

Hybrid
USD 153,000 - 192,000
Discretionary incentive
Benefits package
DevOps Engineer
DevOps Engineer

PowerPlan, Inc. • Atlanta (GA)

Hybrid
USD 120,000 - 155,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000