Manager, Site Reliability Engineering – Paylo Platform

Pditechnologies

Alpharetta (GA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

PDI Technologies is seeking a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel‑pricing product suite. You will manage a ~20‑person team, set reliability roadmaps, and stay hands‑on to maintain high availability across AWS and Azure.

This hands‑on, leadership‑first role requires guiding architects, reviewing designs, and driving CI/CD, IaC, and observability practices with Datadog and Argo CD.

Qualifications

  • 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure/Platform Engineering, including 4+ years in a people‑leadership role.
  • Proven experience managing managers — directly led team leads/managers at scale (~20 reports).
  • Strong hands‑on expertise across AWS and Azure and multi‑cloud infra.
  • Hands‑on expertise with Kubernetes, Helm, and GitOps workflows.
  • Experience with IaC (Terraform/OpenTofu) and CI/CD (Jenkins).
  • Experience with Datadog observability platform and incident management.

Responsibilities

  • Directly manage and develop 3 SRE Managers/Leads and ~20 engineers.
  • Set the reliability roadmap aligning with business priorities.
  • Partner with engineering, product, and stakeholders to balance reliability and risk.
  • Contribute to architecture reviews, troubleshoot incidents, and enforce standards.
  • Own IaC strategy across Terraform/OpenTofu and multi‑cloud platforms.
  • Oversee CI/CD pipelines and deployment strategies (blue/green, canary).
  • Drive incident management, postmortems, and long‑term reliability investments.
  • Drive cost, capacity, and KPI reporting to senior leadership.

Skills

Leadership
AWS
Azure
Kubernetes
Terraform/OpenTofu
Jenkins
Datadog
GitOps
CI/CD
Incident management

Tools

Argo CD
Helm
Jenkins
Datadog

Job description

At PDI Technologies, we empower some of the world's leading convenience retail and petroleum brands with cutting‑edge technology solutions that drive growth and operational efficiency.By “Connecting Convenience” across the globe, we empower businesses to increase productivity, make more informed decisions, and engage faster with customers through loyalty programs, shopper insights, and unmatched real-time market intelligence via mobile applications, such as GasBuddy. We’re a global team committed to excellence, collaboration, and driving real impact. Explore our opportunities and become part of a company that values diversity, integrity, and growth.

Role Overview

PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel‑pricing product suite. This role owns the reliability, infrastructure, and operational strategy for a portfolio of high‑traffic, customer‑and‑partner‑facing platforms that power payment transactions, fuel pricing, loyalty and rewards, and offer/coupon redemption for convenience retail and fuel customers around the world.

This is a hands‑on, leadership‑first role. You will manage a team of three SRE Managers/Leads who together lead approximately 20 engineers, while staying technically engaged yourself — reviewing architecture, unblocking hard infrastructure problems, and setting the technical bar across the organization. You will bring strong, current, hands‑on expertise across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/OpenTofu, Jenkins, and Datadog, and you will be a strong, visible people leader who can coach managers and represent SRE to senior engineering and business stakeholders.

  • Directly manage and develop 3 SRE Managers/Leads and own the overall health, growth, and performance of an ~20‑person SRE organization supporting the Paylo product suite.

  • Set the vision, priorities, and operating cadence for the SRE function; translate business and product priorities into a reliability roadmap your managers can execute against.

  • Build a strong bench by hiring, coaching, and developing managers and senior engineers while creating clear career paths and succession plans.

  • Foster a blameless, learning‑oriented culture around incidents, on‑call, and operational excellence.

  • Partner closely with engineering directors, product managers, and business stakeholders across the Paylo organization to align reliability investments with business risk and customer impact.

  • Stay technically engaged day to day by participating in architecture and design reviews, troubleshooting complex production issues, and directly contributing to infrastructure‑as‑code, Kubernetes manifests/Helm charts, and CI/CD pipelines when needed.

  • Set and enforce engineering standards for multi‑cloud infrastructure across AWS and Azure and for container orchestration on Kubernetes at scale.

  • Own adoption and standards for GitOps‑based continuous delivery using Argo CD/Argo Workflows, including deployment strategy, rollout policy, and multi‑cluster promotion.

  • Own the Infrastructure‑as‑Code strategy across teams (Terraform, OpenTofu), including module standards, state management, drift detection, and remediation.

  • Own CI/CD pipeline architecture and standards built on Jenkins, driving build/deploy automation, pipeline reliability, and progressive delivery practices such as blue‑green/canary deployments and automated rollback.

  • Evaluate and guide adoption of new infrastructure tooling and patterns as the platform evolves across AWS and Azure.

  • Own the observability strategy across all supported products, with deep, hands‑on expertise in Datadog (APM, infrastructure monitoring, log management, dashboards, and alerting) as the standard platform for metrics, tracing, and alerting.

  • Define and drive adoption of SLIs/SLOs, error budgets, and reliability KPIs across the organization, holding managers and teams accountable to them.

  • Own the incident management program end to end, including on‑call structure, escalation paths, severity definitions, postmortems, and follow‑through on remediation actions.

  • Drive root‑cause analysis and long‑term reliability investments that reduce Sev1/Sev2 frequency and recurrence.

  • Ensure appropriate resilience, disaster recovery, and capacity planning practices are in place given the sensitivity of payment‑and‑transaction‑related systems.

  • Partner with Security and Compliance to maintain awareness of PCI DSS and related compliance requirements and ensure the SRE organization supports audit and compliance readiness.

  • Track and report cost, capacity, and operational KPIs to senior leadership.

  • 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure/Platform Engineering, including 4+ years in a people‑leadership role.

  • Proven experience managing managers — you have directly led team leads/managers, not just individual contributors, and are comfortable operating at the scale of ~20 total reports.

  • Strong, hands‑on expertise across AWS and Azure — you can architect, troubleshoot, and operate multi‑cloud infrastructure yourself, not just direct others to do so.

  • Strong, hands‑on expertise with Kubernetes and Helm — cluster operations, troubleshooting at scale, and chart design/maintenance.

  • Strong, hands‑on expertise with Argo CD/Argo Workflows for GitOps‑based continuous delivery.

  • Strong, hands‑on expertise with Infrastructure as Code (Terraform, OpenTofu), including module design and state management.

  • Strong, hands‑on expertise with Jenkins for CI/CD pipeline design, administration, and automation.

  • Strong, hands‑on expertise with Datadog (or equivalent enterprise observability platform), including designing monitoring/alerting strategy, dashboards, and APM/tracing at scale.

  • Demonstrated track record of driving incident management, on‑call, and postmortem programs for high‑traffic, customer‑facing systems.

  • Excellent communication and stakeholder‑management skills; able to represent SRE to engineering leadership and business partners with equal credibility.

  • Applicants must be legally authorized to work in the United States without the need for employer sponsorship, now or in the future. PDI Technologies is unable to offer visa sponsorship for this role.

  • Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations.

  • Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect, Microsoft Certified: Azure Solutions Architect, or HashiCorp Terraform Associate.

  • Experience with messaging systems (Kafka/SQS/SNS), PagerDuty (or similar), and multi‑region/multi‑AZ resilience patterns.

  • Prior experience consolidating or standardizing SRE and DevOps practices across multiple product lines or recently‑integrated/acquired teams.

  • Experience partnering with product and business stakeholders to translate reliability investments into business outcomes.

  • A stable, well‑led SRE organization with clear ownership, career paths, and low regrettable attrition among your managers and their teams.

  • Consistent, Datadog‑driven observability and SLOs in place across the organization, with measurable reduction in Sev1/Sev2 incidents and mean time to detect/resolve.

  • Modern, standardized infrastructure practices — GitOps delivery via Argo, IaC via Terraform/OpenTofu, and reliable CI/CD via Jenkins — adopted consistently across teams and clouds.

  • A mature, blameless incident‑management culture with strong postmortem follow‑through.

  • Strong cross‑functional trust with engineering, product, and security/compliance stakeholders.

PDI is committed to offering a well‑rounded benefits program, designed to support and care for you, and your family throughout your life and career. This includes a competitive salary, market‑competitive benefits, and a quarterly perks program. We encourage a good work‑life balance with ample time off [time away] and, where appropriate, hybrid working arrangements. Employees have access to continuous learning, professional certifications, and leadership development opportunities. Our global culture fosters diversity, inclusion, and values authenticity, trust, curiosity, and diversity of thought, ensuring a supportive environment for all.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Head of SRE - Paylo Platform & Multi-Cloud Reliability
Head of SRE - Paylo Platform & Multi-Cloud Reliability

Pditechnologies • Alpharetta (GA)

On-site
USD 150,000 - 210,000
Sr. Manager, Engineering
Sr. Manager, Engineering

Pditechnologies • Alpharetta (GA)

Hybrid
USD 170,000 - 230,000
Competitive salary
Market-competitive benefits
Quarterly perks program
+2
Software Engineer II
Software Engineer II

Socket.dev • Temple (TX)

Hybrid
USD 85,000 - 110,000
Software Engineer II
Software Engineer II

Lever, Inc. • Alpharetta (GA)

Hybrid
USD 95,000 - 135,000
Hybrid work arrangements
Professional learning benefits
Competitive compensation
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Priority Commerce • United States

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
HSA and FSA options
+6
Software Engineer II
Software Engineer II

Lever, Inc. • Houston (TX)

Hybrid
USD 92,000 - 125,000
Hybrid work arrangements
Quarterly perks program
Continuous learning opportunities
+1
Software Engineer II
Software Engineer II

Lever, Inc. • Temple (TX)

Hybrid
USD 90,000 - 120,000
Hybrid working arrangements
Competitive salary and benefits
Quarterly perks program
Software Engineer III
Software Engineer III

PDI Technologies • Houston (TX)

On-site
USD 110,000 - 140,000
Hybrid work arrangements
Quarterly perks program
Continuous learning opportunities
Software Engineer III
Software Engineer III

PDI Technologies • Alpharetta (GA)

On-site
USD 110,000 - 160,000
Hybrid work
Learning & development
Leadership development
+1
Software Engineer III
Software Engineer III

Socket.dev • Temple (TX)

Hybrid
USD 110,000 - 140,000
Hybrid working arrangements
Quarterly perks program
Learning; certifications and growth