Vice President - Site Reliability Engineering

Goldman Sachs

Birmingham

On-site

GBP 60,000 - 100,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Goldman Sachs is seeking an experienced Site Reliability Engineer to lead the design and operation of highly available, resilient platforms. You will drive SLOs, incident response, and automation across multiple engineering teams.

Ideal candidates have 7–10 years in SRE or related roles, deep cloud expertise (AWS/GCP/Azure), and strong coding and debugging skills in Java, Python, or Node.js. This role emphasizes blameless post-mortems and toil reduction.

Qualifications

  • 7–10 years of relevant experience in SRE or software engineering.
  • Strong proficiency in Java, Python, or Node.js with focus on clean, maintainable code for tooling and automation.
  • Hands-on experience with IaC frameworks and cloud platforms.

Responsibilities

  • Partner with engineering leadership to establish SLIs, SLOs, and error budgets.
  • Collaborate to architect highly available, fault-tolerant systems.
  • Conduct architectural reviews; introduce circuit breakers and rate limiting.
  • Reduce toil by building automation and self-service capabilities.
  • Improve production readiness via load testing and chaos engineering.
  • Lead incident response; drive blameless post-mortems and preventive actions.
  • Promote healthy on-call models with clear escalation paths.
  • Help design observable, resilient platforms for critical services.
  • Collaborate across teams to improve production architecture and delivery.

Skills

Java
Python
Node.js
Analytical thinking
SRE mindset

Education

Bachelor's degree in Computer Science or related

Tools

Terraform
Ansible
CloudFormation
Docker
Kubernetes
AWS
GCP
Azure
Prometheus
Grafana
Datadog
OpenTelemetry
ELK
CloudWatch

Job description

Salary: £60,000 - 100,000 per year

Requirements
  • Strong proficiency in at least one major programming language such as Java, Python, or Node.js, with a focus on clean, maintainable code for tooling and automation.
  • Hands-on experience with Infrastructure as Code frameworks such as Terraform, Ansible, or CloudFormation.
  • Deep understanding of containerization and orchestration technologies, specifically Docker and Kubernetes, including service meshes and ingress controllers.
  • Advanced experience with major cloud providers such as AWS, GCP, or Azure, building and operating highly resilient cloud-native architectures.
  • Proficiency with observability stacks, including distributed tracing, logging, and metrics such as Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch.
  • Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures, and software design.
  • Knowledge of networking protocols, VPCs, load balancing strategies in distributed systems, and database query performance tuning.
  • Ability to analyze complex distributed systems holistically and understand how individual components interact under load.
  • Strong interpersonal skills to collaborate with product developers, influence architectural decisions, prioritize toil reduction, and drive SRE adoption without direct authority.
  • Ability to translate complex technical issues into clear, actionable insights for both technical and non-technical stakeholders.
  • Highly motivated, proactive, and capable of multitasking under pressure in a fast-paced environment without compromising quality.
  • Commitment to fostering a blameless culture where failures are treated as opportunities to learn and improve systems.
  • Interest in financial markets and technology.
  • Bachelors degree in Computer Science, System Engineering, or a related technical field that involves programming.
  • 7 to 10 years of experience.
Responsibilities
  • Partner with engineering leadership to establish service level objectives, service level indicators, and error budgets.
  • Collaborate with product developers to architect highly available, fault-tolerant, and self-healing systems.
  • Conduct architectural reviews and introduce patterns like circuit breakers, graceful degradation, and rate limiting.
  • Reduce operational toil by building automation, tooling, and self-service capabilities that remove repetitive manual work.
  • Improve production readiness through load testing, performance tuning, capacity forecasting, chaos engineering, and reliability reviews.
  • Lead the response to complex, multi-system production incidents.
  • Facilitate blameless post-mortems to identify root causes and drive long-term preventative actions.
  • Promote sustainable operations by helping design healthy on-call models, clear escalation paths, and balanced pager responsibilities.
  • Help engineer highly reliable, observable, and resilient platforms that support critical business services at scale.
  • Collaborate with multiple engineering teams to continually improve production system architecture, facilitate fast delivery of new services, and reduce downtime.
  • Champion SRE principles such as SLOs, error budgets, and blameless post-mortems across a large engineering organization.
Technologies
  • AWS
  • Ansible
  • Architect
  • Azure
  • Cloud
  • CloudWatch
  • Datadog
  • Docker
  • ELK
  • GCP
  • Grafana
  • Support
  • Java
  • Kubernetes
  • Linux
  • Load Balancing
  • OpenTelemetry
  • Prometheus
  • Python
  • Splunk
  • Terraform
  • NodeJS
More

We are Goldman Sachs, a leading global investment banking, securities, and investment management firm founded in 1869 and headquartered in New York, with offices in all major financial centers around the world. Our Site Reliability Engineering team sits at the intersection of software engineering, systems design, and production excellence, building highly reliable, observable, and resilient platforms that support critical business services at scale. We commit our people, capital, and ideas to help our clients, shareholders, and communities grow, and we are committed to diversity, inclusion, professional development, wellness, personal finance offerings, and mindfulness programs. We also support candidates with special needs or disabilities during our recruiting process.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham
Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham

Worky • Birmingham

On-site
GBP 150,000 - 210,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JP Morgan Chase • Glasgow

On-site
GBP 62,000 - 102,000
Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President
Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Goldman Sachs Bank AG • Birmingham

Hybrid
GBP 120,000 - 180,000
Healthcare & Medical Insurance
Financial Wellness & Retirement
On-site health centers
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1
Head of Site Reliability Engineering
Head of Site Reliability Engineering

Goldman Sachs • Birmingham

On-site
GBP 60,000 - 100,000
The Core Engineering - Site Reliability Engineering - Associate - Birmingham
The Core Engineering - Site Reliability Engineering - Associate - Birmingham

WeAreTechWomen • Birmingham

On-site
GBP 65,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Daily catered lunches
Modern office environment
Tech talks and knowledge sharing
SRE
SRE

Source Group International • Greater London

Hybrid
GBP 68,000 - 108,000
IAM Secrets Management Engineering - SRE Platform Engineer - VP - London London · United Kingdo[...]
IAM Secrets Management Engineering - SRE Platform Engineer - VP - London London · United Kingdo[...]

Goldman Sachs Bank AG • Greater London

On-site
GBP 80,000 - 110,000