Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work
On-call rotation

Job summary

T. Rowe Price is seeking a Principal Site Reliability Engineer, Infrastructure Observability, to lead a team of SREs focused on observability, sustainability, scalability, measurability and recoverability of cloud and on‑prem solutions.

The role emphasizes automation, incident response, and 24x7 monitoring within a complex, distributed environment. The successful candidate will have a hands‑on operations and engineering background, deep cloud experience (public/private), and expertise in CI/CD,

Qualifications

  • Bachelor’s degree or equivalent combination of education and relevant experience with 10+ years in cloud infrastructure design/operation.
  • 5+ years building and supporting AWS-based solutions.
  • 5+ years building/running a DevOps and/or SRE function.
  • Experience with chaos engineering at scale.
  • Proven track record implementing new technology, tools, and platforms.

Responsibilities

  • Design technology solutions to prevent or minimize service disruptions.
  • Lead automation initiatives to reduce outages and improve reliability across distributed systems.
  • Drive blameless post-mortems and learning to improve reliability across services.
  • Transform operations teams to adopt SRE practices and drive strategic growth.
  • Analyze incidents to identify trends and prevent recurrence.
  • Define target state architecture and portfolio-wide observability standards.
  • Collaborate with multiple partners and sponsor groups to align on roadmaps.

Skills

Cloud infrastructure
AWS
DevOps/SRE
Automation
Incident response
Observability
SRE governance
Programming (Python/Java)

Education

Bachelor’s degree or equivalent experience

Tools

New Relic
Elastic Stack
Prometheus
Grafana
Splunk
Terraform
Ansible

Job description

KM5

Role Summary

In this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the observability, sustainability, scalability, measurability and recoverability of T. Rowe Price’s innovative cloud & on-prem solutions by leveraging automation and best-of-breed tools. The successful candidate will have a strong operations & engineering background, is hands-on when needed, and has expertise in the cloud environments (public, private), infrastructure operations, DevOps practices, CI/CD toolchain and systems, code build and deployment, incident response, and 24x7 monitoring and support.

The candidate will also have extensive experience operating within a SRE function within a complex, distributed environment. They will have a demonstrated ability to work horizontally and vertically within an organization with diverse partners and sponsor groups.

Responsibilities

  • Possesses extensive knowledge in own area of expertise and extensive in-depth knowledge of the broader portfolio for comprehensive understanding of up/downstream impacts across technology infrastructure
  • Responsibility for the design of technology solutions to prevent or minimize service disruptions
  • Prevents technology service disruptions through technology solution recommendations and automations
  • Fosters a culture of deep learning through blameless post‑mortems to improve the shared goal of reliability across services
  • Transform operations teams by facilitating internal change to adopt SRE standard methodologies across the organization and driving strategic growth in this area within Global Technology
  • Analyzes incidents impacting technology availability for high‑level trends across the broad portfolio
  • Drive initiatives to reduce or prevent technology failures in a complex, distributed technology environment
  • Pulls together information from disconnected systems into cohesive views of the technology portfolio for identifying trends, redundancies, and risk
  • Demonstrates outstanding awareness of the complexities of the tech and asset management industries
  • May lead initiatives of varying degrees of complexity that span multi‑functional areas and of varying degrees of complexity
  • Contributes to definition of target state architecture and design of the technology environment

Qualifications
Required:

  • Bachelor’s degree or the equivalent combination of education and relevant experience AND 10+ years of experience designing and operating cloud infrastructure with senior‑level impact.
  • 5+ years building and supporting solutions in Amazon AWS
  • 5+ years of experience building and running a DevOps and/or SRE function
  • Experience with implementation and operation of the chaos model at scale
  • Strategic and program‑level implementation experience
  • Demonstrable experience implementing new technology, tools, and platforms
  • System administration and scripting experience
  • Demonstrable experience leveraging automation to proactively prevent or quickly remediate incidents
  • Fluent in multiple programming languages (e.g., Python, Java, GO, Node.js, .Net Core, etc)
  • Proficiency with database development (SQL Server, PostgreSQL, MySQL, etc)
  • Proficiency with defining, right‑size‑ing, tracking, and reporting on Service Level Objectives (SLOs), Service Level Indicators (SLIs), system availability, and the progress and outcomes related to reliability
  • Experience with implementing and managing Error Budgets
  • Proficiency with understanding and explaining incident situations and their recovery plans to prevent recurrence
  • Knowledge/experience driving dashboard standardization across the ecosystem for observability, APM and infrastructure monitoring, and application‑specific logging
  • Knowledge/experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk, and cloud native tools
  • Knowledge/experience with cloud management tools such as Ansible, Terraform, Vault, and Vagrant
  • Works independently, with guidance in only the most complex situations
  • Makes sound decisions with limited facts or resources
  • Balances strategic and pragmatic concerns when solving problems
  • Adjusts communication style and materials to suit a given audience
  • Able to clearly articulate operational principles, practices, and policies
  • Stays abreast of industry trends and technologies
  • Accountable for work of self and others; sets standards around which others will operate
  • Maintains a broad internal professional network and knows when to engage/activate it
  • Develops or mentor’s diverse talent on the team
  • Ability to be on‑call and/or work during off‑hours

Preferred:

  • Cloud or SRE‑related certifications
  • Working knowledge of Azure

**Work Flexibility**
This role is eligible for hybrid work, with up to three days per week from home.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Senior SRE: Observability & Cloud Reliability
Senior SRE: Observability & Cloud Reliability

T. Rowe Price • Greater London

Hybrid
GBP 120,000 - 180,000
Hybrid work
On-call rotation
Senior Site Reliability Engineer - Selby Jennings
Senior Site Reliability Engineer - Selby Jennings

eFinancialCareers • Greater London

On-site
GBP 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tenth Revolution Group • Knutsford

Hybrid
GBP 70,000 - 90,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Senior Site Reliability Engineer — Global Cloud & Automation
Senior Site Reliability Engineer — Global Cloud & Automation

Boston Consulting Group (BCG) • Greater London

Hybrid
GBP 70,000 - 90,000
Senior SRE Architect & Platform Reliability Lead
Senior SRE Architect & Platform Reliability Lead

Boston Consulting Group (BCG) • Greater London

Hybrid
GBP 80,000 - 120,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Boston Consulting Group (BCG) • Greater London

On-site
GBP 80,000 - 120,000
Observability Engineer - Assistant Vice President
Observability Engineer - Assistant Vice President

Citibank (Switzerland) AG • Greater London

Hybrid
Confidential
Annual leave 27d
Discretionary bonus
Medical & life insurance
+5
Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

Hybrid
GBP 51,000 - 85,000
Bonus
Benefits