VP Site Reliability Engineer — Global Markets

Goldman Sachs

New York (NY)

On-site

USD 150,000 - 250,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Goldman Sachs is seeking an experienced Site Reliability Engineer to design and operate high-availability, multi-region cloud-native services, with strong emphasis on reliability, observability, and AI-assisted automation.

You will define SLOs/SLIs, manage incident responses, and drive modernization across production systems in a fast-paced financial environment. The role requires 8+ years in software/reliability engineering and deep cloud/Kubernetes experience.

Qualifications

  • 8+ years of professional software/reliability engineering experience.
  • Proficiency in Java 17+ and concurrency.
  • Demonstrated risk acumen in regulated financial services.
  • Excellent communication and stakeholder coordination.
  • Experience running high-availability production environments with SLIs/SLOs and on-call.
  • Strong cloud infrastructure knowledge (GCP, AWS) and container orchestration (Kubernetes, Docker).
  • Experience with AI models/tools governance and QA of AI-generated work.
  • Experience building event-driven and distributed systems (Kafka) and CI/CD.

Responsibilities

  • Design, build, and operate high-availability, multi-region, cloud-native services with security and observability built in at every layer.
  • Establish and manage SLIs, SLOs, and error budgets; drive blameless post-incident reviews and translate findings into durable engineering improvements.
  • Lead incident response for latency-sensitive, high-throughput trade lifecycle systems — quickly diagnosing issues, coordinating cross-functional responders, and communicating status to stakeholders.
  • Develop event-driven architectures, multi-stage processing pipelines, and optimized data paths for high-throughput trade lifecycle management.
  • Apply strong risk acumen to change management, capacity planning, and resilience testing (chaos engineering, failover, and BCP drills).
  • Partner with engineers, domain experts, and global stakeholders to understand production processes, challenge entrenched assumptions in a cloud-centric, AI-driven world, and drive modernization.
  • Multiply your impact with a modern, AI-centric toolchain, orchestrating AI agents across the SDLC and operations to rapidly comprehend large codebases, generate production-quality automation, and accelerate delivery.

Skills

Java 17+
Concurrency
Cloud experience
Stakeholder coordination
Incident management
AI tooling awareness

Tools

Kubernetes
Docker
Terraform
GitHub Copilot
Claude Code
Apache Kafka
Prometheus
OpenTelemetry
Spring Boot

Job description

Goldman Sachs is seeking an experienced Site Reliability Engineer to design and operate high-availability, multi-region cloud-native services, with strong emphasis on reliability, observability, and AI-assisted automation.

You will define SLOs/SLIs, manage incident responses, and drive modernization across production systems in a fast-paced financial environment. The role requires 8+ years in software/reliability engineering and deep cloud/Kubernetes experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VP Site Reliability Engineer: AI-Driven Global Markets
VP Site Reliability Engineer: AI-Driven Global Markets

The Goldman Sachs Group • New York (NY)

On-site
USD 150,000 - 250,000
Senior SRE - Global Markets & AI Ops
Senior SRE - Global Markets & AI Ops

Socket.dev • New York (NY)

On-site
USD 150,000 - 300,000
SRE VP, Global Banking & Markets: High-Impact Reliability
SRE VP, Global Banking & Markets: High-Impact Reliability

Goldman Sachs Bank AG • New York (NY)

On-site
USD 150,000 - 250,000
VP, SRE Platforms — Scale, Reliability & Automation
VP, SRE Platforms — Scale, Reliability & Automation

The Goldman Sachs Group • Dallas (TX)

On-site
USD 180,000 - 280,000
Platform SRE Engineer — Reliability & Observability
Platform SRE Engineer — Reliability & Observability

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
None
SRE Platforms Engineer: Reliability, Observability & Scale
SRE Platforms Engineer: Reliability, Observability & Scale

The Goldman Sachs Group • Dallas (TX)

On-site
USD 110,000 - 140,000
SRE Platform Engineer: Reliability & Observability
SRE Platform Engineer: Reliability & Observability

Goldman Sachs • Dallas (WV)

On-site
USD 120,000 - 180,000
VP, SRE Platforms & Reliability
VP, SRE Platforms & Reliability

Goldman Sachs, Inc. • Dallas (TX)

On-site
USD 210,000 - 260,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
None
Engineering – SRE Platforms – Software Engineer – Vice President – Dallas | Dallas, TX, USA
Engineering – SRE Platforms – Software Engineer – Vice President – Dallas | Dallas, TX, USA

Goldman Sachs, Inc. • Dallas (TX)

On-site
USD 210,000 - 260,000