Senior Site Reliability Engineer

Vanguard

Charlotte (NC)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hybrid work model

Job summary

Vanguard is seeking a Site Reliability Engineer (SRE) to lead resiliency efforts within Global Technology Operations. You will design, test, and strengthen critical systems, automate incident response, and push AI‑enhanced diagnostics to improve detection, response, and recovery.

You will collaborate with cross‑functional teams to raise operational maturity and deliver reliable, client‑centric experiences as Vanguard continues its hybrid work model.

Qualifications

  • Experience with observability, monitoring, and reliability metrics.
  • Strong understanding of SLIs, SLOs, and SLAs.
  • Expertise in alert design and anomaly detection.
  • Hands-on automation and resilience engineering capabilities.

Responsibilities

  • Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.
  • Design and implement processes that enforce enterprise resiliency and reliability standards.
  • Lead blameless post‑incident reviews for high‑severity incidents or incidents spanning multiple complex product families.
  • Partner with product and platform teams to proactively identify and remediate reliability risks before they impact clients.
  • Develop, communicate, and evangelize new standards, tools, and frameworks across subdivisions, ensuring consistent adoption.
  • Troubleshoot complex production issues and implement durable solutions that prevent recurrence.
  • Participate in a periodic on‑call rotation to support production stability.
  • Evaluate and onboard resiliency and reliability tooling.
  • Actively participate in reliability engineering and resilience communities of practice, contributing to shared learning and enterprise consistency.
  • Contribute to strategic initiatives that advance Vanguard’s operational maturity and resiliency posture.

Skills

Observability platforms
SLIs/SLOs/SLAs
Monitoring & alerting
Automation & resilience engineering

Tools

Splunk
Honeycomb
CloudWatch
Dynatrace
AppDynamics

Job description

The Site Reliability Engineer (SRE) for Global Technology Operations (GTO) is a strategic technical leader responsible for ensuring the resiliency and stability of the critical applications our crew and clients rely on every day. This role combines deep hands‑on engineering expertise with enterprise‑level influence. You will help define what resiliency means at Vanguard and partner across teams to design, test, and strengthen some of our most critical systems. In addition, you will automate incident response capabilities and pioneer AI‑enhanced diagnostics and analysis to improve detection, response, and recovery. You will work alongside a collaborative, technically focused team where your innovations in resiliency engineering directly shape Vanguard’s next generation of reliable, client‑centric experiences.

Core Responsibilities
  • Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.
  • Design and implement processes that enforce enterprise resiliency and reliability standards.
  • Lead blameless post‑incident reviews for high‑severity incidents or incidents spanning multiple complex product families.
  • Partner with product and platform teams to proactively identify and remediate reliability risks before they impact clients.
  • Develop, communicate, and evangelize new standards, tools, and frameworks across subdivisions, ensuring consistent adoption.
  • Troubleshoot complex production issues and implement durable solutions that prevent recurrence.
  • Participate in a periodic on‑call rotation to support production stability.
  • Evaluate and onboard resiliency and reliability tooling.
  • Actively participate in reliability engineering and resilience communities of practice, contributing to shared learning and enterprise consistency.
  • Contribute to strategic initiatives that advance Vanguard’s operational maturity and resiliency posture.
Qualifications | Technical Skills
  • Observability Platforms: Experience with modern observability and monitoring tools, such as Splunk, Honeycomb, CloudWatch, Dynatrace, or AppDynamics.
  • Reliability Metrics: Strong understanding of SLIs, SLOs, and SLAs, including dashboarding and reporting practices.
  • Monitoring & Alerting: Experience with alert design, anomaly detection, predictive alerting, and synthetic monitoring using structured methodologies.
  • Automation & Resilience Engineering: Experience with automation and resilience practices such as Python-based automation, RPA platforms (e.g., Blue Prism, UiPath), chaos engineering, and failure analysis techniques (e.g., FMEA).
Special Factors
Sponsorship

Vanguard is not offering visa sponsorship for this position.

About Vanguard

At Vanguard, we don't just have a mission—we're on a mission.

To work for the long-term financial wellbeing of our clients.

To lead through product and services that transform our clients' lives.

To learn and develop our skills as individuals and as a team.

From Malvern to Melbourne, our mission drives us forward and inspires us to be our best.

How We Work

Vanguard has implemented a hybrid working model for the majority of our crew members, designed to capture the benefits of enhanced flexibility while enabling in‑person learning, collaboration, and connection.

We believe our mission‑driven and highly collaborative culture is a critical enabler to support long‑term client outcomes and enrich the employee experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineer, Site Reliability
Engineer, Site Reliability

Vanguard • Malvern

Hybrid
USD 150,000 - 230,000
Full Stack Software Engineer, Resiliency Engineering Platforms
Full Stack Software Engineer, Resiliency Engineering Platforms

The Vanguard Group • Wayne (PA)

Hybrid
USD 110,000 - 170,000
Hybrid work model
Full Stack Software Engineer, Resiliency Engineering Platforms
Full Stack Software Engineer, Resiliency Engineering Platforms

Vanguard • North Carolina

Hybrid
USD 110,000 - 160,000
Hybrid work model
Engineering Manager, Site Reliability
Engineering Manager, Site Reliability

Vanguard • Malvern

Hybrid
USD 180,000 - 240,000
Senior SRE: Enterprise Resiliency & AI-Driven Reliability
Senior SRE: Enterprise Resiliency & AI-Driven Reliability

Vanguard • Charlotte (NC)

Hybrid
USD 120,000 - 160,000
Hybrid work model
Cloud Security Engineer, Specialist
Cloud Security Engineer, Specialist

Vanguard • Malvern

Hybrid
USD 150,000 - 190,000
Security Engineering Governance & Data Analyst
Security Engineering Governance & Data Analyst

The Vanguard Group • Malvern

Hybrid
USD 110,000 - 170,000
Hybrid work model
Manager, IT Delivery
Manager, IT Delivery

The Vanguard Group • East Whiteland Township (PA)

Hybrid
USD 150,000 - 210,000
Technology Leadership Program, Risk and Security Engineer (TX)
Technology Leadership Program, Risk and Security Engineer (TX)

Vanguard • Dallas (TX)

Hybrid
USD 85,000 - 110,000
Hybrid work model
Competitive compensation
Healthcare benefits
+3
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 180,000 - 280,000