Principal Site Reliability Engineer

Gen

Tempe (AZ)

On-site

USD 150,000 - 160,000

Full time

30 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Gen Digital is seeking a Principal Site Reliability Engineer to define reliability strategy and lead cross‑functional SRE initiatives in Tempe, AZ. You will own production readiness, automation, and incident response for a cloud‑native platform.

You will guide architecture, establish SRE practices, and collaborate with security to meet PCI/SOC 2‑level requirements while mentoring senior engineers and driving global best practices.

Qualifications

  • 8+ years in SRE/DevOps or infrastructure engineering leading large cloud projects.
  • Strong expertise in AWS and Kubernetes for large-scale microservices.
  • Experience with Terraform and infrastructure as code.
  • Proven ability to define SRE practices such as SLIs/SLOs, incident management, observability.
  • Excellent written and verbal communication with senior stakeholders.

Responsibilities

  • Define and drive long‑term reliability strategy across services and teams.
  • Lead architecture decisions for highly available distributed systems on AWS, Kubernetes, and cloud platforms.
  • Establish SRE practices including SLIs, SLOs, error budgets, and production readiness reviews.
  • Guide platform architecture and automation using infrastructure‑as‑code tooling.
  • Advance observability with metrics, logging, tracing, and dashboards.
  • Lead major incident responses and drive durable post‑incident actions.
  • Mentor senior engineers and align regional SRE practices with global standards.
  • Champion resilience, disaster recovery, and capacity planning for mission‑critical systems.
  • Collaborate with security/compliance to strengthen controls in regulated environments.
  • Improve CI/CD reliability and developer experience with modern tooling.

Skills

SRE leadership
Cloud architecture
Incident management
Observability
Automation
CI/CD
Python or Go
Security/compliance
Communication
Mentoring

Tools

AWS
Kubernetes
GKE
Terraform
Jenkins
CircleCI
GitHub Actions

Job description

About The Role

At Gen Digital your role will be supporting products and services, including Gen Insurance and Engine by Gen, delivering trusted digital and financial experiences to consumers at scale.

About The Role

At Gen Digital your role will be supporting products and services, including Gen Insurance and Engine by Gen, delivering trusted digital and financial experiences to consumers at scale.

As a Principal Site Reliability Engineer, you will define and lead the reliability, scalability, and operational excellence strategy for our platform. This is a senior technical leadership role for an engineer who combines deep hands‑on expertise with broad organizational influence. You will partner across application, platform, security, and infrastructure teams to build resilient systems, modernize our cloud‑native architecture, and raise engineering standards across the company.

You will be responsible for shaping the next stage of our SRE practice, including service reliability frameworks, production readiness, observability, incident management, automation, and capacity planning. You will also collaborate closely with other SRE teams to align on global standards, shared practices, and strategic platform direction.

Key Responsibilities
  • Define and drive the long‑term reliability strategy for the platform, establishing technical direction and operational standards across services and teams.
  • Lead architecture decisions for highly available, fault‑tolerant distributed systems running on AWS, GCP, Kubernetes, and GKE, with a strong focus on scalability, resilience, and security.
  • Establish and mature SRE practices including SLIs, SLOs, error budgets, production readiness reviews, service ownership standards, and operational risk management.
  • Partner with engineering leadership to guide the evolution of platform architecture, deployment patterns, and infrastructure automation using Terraform and other infrastructure‑as‑code tooling.
  • Drive the design and implementation of next‑generation observability capabilities, including metrics, logging, tracing, alerting strategy, and actionable dashboards.
  • Lead major incident response for critical production events, improve escalation paths and response processes, and ensure durable corrective actions through strong post‑incident reviews.
  • Influence the reliability and performance posture of all engineering teams by reviewing RFCs, setting engineering guardrails, and providing technical leadership on high‑impact initiatives.
  • Champion resilience engineering, disaster recovery, failover design, capacity forecasting, and business continuity planning for mission‑critical systems.
  • Partner with security and compliance stakeholders to strengthen infrastructure security and operational controls in regulated and compliant environments, including PCI and SOC 2 requirements.
  • Drive continuous improvement of CI/CD systems and software delivery workflows to improve deployment safety, speed, repeatability, and developer experience.
  • Identify opportunities to reduce operational toil through automation, self‑service platform capabilities, and better engineering abstractions.
  • Mentor senior engineers and technical leads, raising the bar for system design, operational excellence, and production ownership across the organization.
  • Align with global SRE counterparts to standardize best practices, share learnings, and improve collaboration across regions and organizations.
  • Serve as a trusted technical advisor to engineering and product leadership on reliability tradeoffs, infrastructure investments, and operational risk.
About You
  • 8+ years of experience in SRE, DevOps, platform engineering, or infrastructure engineering, with a strong record of leading large‑scale cloud initiatives in production environments.
  • Deep expertise in AWS and Kubernetes, including hands‑on experience designing, operating, and evolving large‑scale containerized micro‑service based systems.
  • Strong experience with infrastructure as code, especially Terraform, and with AWS services supporting distributed systems and container platforms such as EKS, ECS, or GKE.
  • Proven success defining and implementing SRE practices such as SLIs, SLOs, error budgets, incident management, observability, and production readiness standards.
  • Strong understanding of distributed systems reliability, performance, scaling, availability engineering, and failure‑mode analysis.
  • Experience leading complex cross‑team technical initiatives and influencing architecture, standards, and engineering practices beyond direct reporting lines.
  • Strong background in CI/CD and delivery engineering, with experience improving pipeline reliability, deployment safety, and release automation using tools such as Jenkins, CircleCI, and GitHub Actions.
  • Proficiency in at least one programming language such as Python or Go, with the ability to build automation, tooling, and maintainable production‑quality code.
  • Experience operating in security‑conscious and compliant environments, with familiarity in frameworks such as PCI, SOC 2, and NIST.
  • Excellent written and verbal communication skills, with the ability to guide technical decision‑making from senior engineers through executive stakeholders.
  • Experience using AI‑assisted and agentic engineering tools to improve productivity, automation, operational insight, and developer workflows, with sound judgment around quality, security, and reliability.
  • A systems thinker who can balance strategic direction with hands‑on execution and who thrives in high‑scale, high‑ownership environments.
Interview Process
  • Recruiter Screen - Phone Call
  • Technical Interview with Team member - Zoom
  • Hiring Team/Manager - Onsite 2 hours

Compensation Range: $150K - $160K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Apply • Tempe (AZ), Northern (KY)

Hybrid
USD 180,000 - 240,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Gen Digital Inc. • Tempe (AZ)

On-site
USD 180,000 - 230,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

AVG • Tempe (AZ), Northern (KY)

Hybrid
USD 140,000 - 190,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Site Reliability Engineering Team Lead (Principal SRE)
Site Reliability Engineering Team Lead (Principal SRE)

Jobgether • United States

Hybrid
USD 170,000 - 210,000
Annual bonus
Medical, dental, and vision insurance
Life and disability insurance
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Supio • San Francisco (CA)

On-site
USD 170,000 - 220,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000