Engineer, Site Reliability

Vanguard

Malvern (Chester County)

Hybrid

USD 150,000 - 230,000

Full time

22 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Vanguard is seeking a Staff Reliability Engineer for the Personal Wealth Technology (PWTech) Reliability Engineering team in Malvern, PA. The role focuses on architecting and building enterprise-scale resiliency solutions and automating incident responses in a microservices production environment.

You will lead automated diagnostics, distributed tracing at scale, and AI-enhanced analysis. The position requires deep expertise in distributed systems, Java/JS, AWS, and OpenTelemetry, with a strong

Qualifications

  • Eight+ years of related experience with at least two years in development.
  • Undergraduate degree or equivalent; graduate degree preferred.

Responsibilities

  • Lead the strategy, architecture, and evolution of PWTech reliability platforms across hundreds of applications.
  • Design and build production-grade platforms that improve reliability, including automated incident detection and remediation.
  • Drive enterprise observability and diagnostics across cloud-native technologies and stacks.
  • Define resilient patterns (graceful degradation, circuit breakers, failover) and promote adoption.
  • Influence teams and leaders to embed reliability from inception with standards.
  • Lead resolution of complex reliability and production challenges and perform root cause analysis.
  • Participate in special projects as assigned.

Skills

Reliability design
Distributed systems
Java/JS
AWS
Observability

Education

Bachelor's degree
Graduate degree preferred

Tools

OpenTelemetry
Docker
Kubernetes
Python

Job description

Join the Personal Wealth Technology (PWTech) Reliability Engineering team and lead cutting-edge Reliability Engineering initiatives that impact hundreds of applications and millions of investors. You'll architect and build enterprise-scale resiliency solutions, driving our ambitious roadmap. This is an opportunity to combine deep technical expertise with strategic influence —automating incident responses, implementing distributed tracing at scale, and pioneering AI-enhanced diagnostics and analysis. Work alongside a collaborative, technically-focused team where your innovation in resilience engineering will shape Vanguard's next generation of client experiences. At Vanguard, we pride ourselves on delivering an exceptional client experience to all investors; at the core of this experience are systems that reside in a technically complex and constantly evolving resiliency landscape. Passionate, technically skilled engineers are at the center of our resiliency operations, and we are looking to grow our team. We are seeking an experienced engineer with broad, end-to-end software development experience, including operating applications in a microservices environment in production at scale. This role goes beyond feature implementation - it requires someone who can design, build, and support resilient systems from the ground up. As a Staff Reliability Engineer at Vanguard, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance innovative thinking with pragmatism and understand the long-term impacts of technical decisions. You communicate complex ideas clearly and collaborate effectively to deliver scalable solutions.

Core Responsibilities
  • Lead the technical strategy, architecture, and evolution of PWTech reliability engineering platforms and capabilities, ensuring they scale across hundreds of applications and critical client-facing systems.
  • Design and build production-grade software and platforms that improve reliability outcomes, including automated incident detection, diagnostics, remediation, resiliency engineering, and operational intelligence.
  • Drive enterprise observability and diagnostics capabilities, enabling consistent telemetry, distributed tracing, metrics, and operational insights across cloud-native technologies and application stacks.
  • Define and codify resilient application and platform patterns, such as graceful degradation, circuit breakers, load shedding, fault isolation, failover, and automated recovery, driving adoption through reusable software, frameworks, and engineering standards.
  • Influence engineering teams and technology leaders across the organization, establishing technical standards and ensuring reliability is designed into systems from inception.
  • Lead the resolution of complex reliability and production challenges, identifying systemic risks, driving root cause analysis, and engineering durable solutions that improve long-term resilience.
  • Participates in special projects and performs other duties as assigned.
Qualifications
  • Minimum of eight years related experience, with at least two years of development experience.
  • Undergraduate degree or equivalent combination of training and experience. Graduate degree preferred.
Preferred Skills
  • Experience designing, building, and operating production-facing platforms or engineering capabilities that achieve broad adoption and deliver measurable reliability, operational, or business outcomes.
  • Deep expertise in distributed systems architecture, including scalability, availability, resiliency, fault tolerance, performance optimization, and production operations at scale.
  • Strong technical leadership and influence skills, with a demonstrated ability to drive architecture decisions, establish technical standards, and align multiple teams on engineering direction.
  • Deep expertise in Java or JavaScript, with hands-on experience developing and operating software in modern cloud-native and microservices environments.
  • Demonstrated ability to diagnose and resolve complex production issues, perform root cause analysis, and engineer durable solutions that prevent recurrence.
  • Hands-on experience with AWS and modern cloud architecture patterns.
  • Experience with observability and telemetry platforms, including metrics, logging, distributed tracing, and production diagnostics. Experience with OpenTelemetry is strongly preferred.
  • Proficiency with Python or similar scripting languages to develop automation, tooling, and operational workflows.
  • Strong software engineering fundamentals, systems thinking skills, experience solving complex reliability challenges, and the ability to influence across teams and drive engineering best practices
Sponsorship

Vanguard is not offering visa sponsorship for this position.

About Vanguard

At Vanguard, we don't just have a mission—we're on a mission.

To work for the long-term financial wellbeing of our clients. To lead through product and services that transform our clients' lives. To learn and develop our skills as individuals and as a team. From Malvern to Melbourne, our mission drives us forward and inspires us to be our best.

How We Work

Vanguard has implemented a hybrid working model for the majority of our crew members, designed to capture the benefits of enhanced flexibility while enabling in-person learning, collaboration, and connection. We believe our mission-driven and highly collaborative culture is a critical enabler to support long-term client outcomes and enrich the employee experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, Site Reliability
Engineering Manager, Site Reliability

Vanguard • Malvern

Hybrid
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Vanguard • Charlotte (NC)

Hybrid
USD 120,000 - 160,000
Hybrid work model
Full Stack Software Engineer, Resiliency Engineering Platforms
Full Stack Software Engineer, Resiliency Engineering Platforms

The Vanguard Group • Wayne (PA)

Hybrid
USD 110,000 - 170,000
Hybrid work model
Cloud Engineer, Specialist, Resiliency Engineering Platforms
Cloud Engineer, Specialist, Resiliency Engineering Platforms

The Vanguard Group • Wayne (PA)

Hybrid
USD 100,000 - 140,000
Visa sponsorship
Cloud Engineer, Specialist, Resiliency Engineering Platforms
Cloud Engineer, Specialist, Resiliency Engineering Platforms

The Vanguard Group • North Carolina

Hybrid
USD 120,000 - 160,000
Visa sponsorship
Manager, IT Delivery
Manager, IT Delivery

The Vanguard Group • Malvern

Hybrid
USD 180,000 - 240,000
Manager, IT Delivery
Manager, IT Delivery

The Vanguard Group • East Whiteland Township (PA)

Hybrid
USD 150,000 - 210,000
Direct Indexing Cloud Platform Infrastructure & Dev Ops Engineer
Direct Indexing Cloud Platform Infrastructure & Dev Ops Engineer

The Vanguard Group • United States

Hybrid
USD 140,000 - 220,000
Direct Indexing Cloud Platform Infrastructure & Dev Ops Engineer
Direct Indexing Cloud Platform Infrastructure & Dev Ops Engineer

Vanguard • Oakland (CA)

Hybrid
USD 180,000 - 240,000
Head of Advice Portfolio Construction Technology
Head of Advice Portfolio Construction Technology

The Vanguard Group • East Whiteland Township (PA)

Hybrid
USD 230,000 - 380,000