Senior SRE: Incident Commander, Automation & Observability

Goldman Sachs Bank AG

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Goldman Sachs is seeking a Senior Site Reliability Engineer (SRE) specializing in incident management, escalation discipline, and automation. You will partner with front-office desks, engineering, and ABO to improve desk readiness and reduce toil through automation and AI.

You will lead cross-region handoffs, governance of error budgets, and establish clear SRE KPIs. The role emphasizes observability, scalable systems, and reliable incident communications across globally distributed

Qualifications

  • Min. 5 years in SRE, production operations, or reliability-focused engineering.
  • Proven experience as Incident Commander with improvements in escalation timeliness and MTTR.
  • Strong foundations in Linux, networking (DNS, HTTP, TLS, routing), distributed systems, and public cloud (AWS/Azure/GCP).
  • Hands-on with observability stacks (Prometheus, Grafana, OpenTelemetry, ELK), incident tooling (PagerDuty, Opsgenie), and collaboration platforms (Slack/Teams).
  • Proficiency with infrastructure-as-code and automation (Terraform, CloudFormation, Ansible) and at least one modern programming language (Go, Python).
  • Experience implementing SLO/SLI/error budgets, capacity planning, progressive delivery, and chaos/game days.
  • Excellent written and verbal communication; able to translate complex technical contexts into concise updates for executives and business stakeholders.
  • Comfortable working across time zones with strong ownership of cross-region handoffs and follow-through.

Responsibilities

  • Incident Command, Escalation, and Communications: Act as Incident Commander for high-severity events with timely escalation and transparent communications to stakeholders.
  • Cross-Region Handoffs and Desk Readiness: Own cross-region handoff procedures with explicit ownership and desk-readiness checklists.
  • ABO Partnership and Workload Reduction: Partner with ABO to identify incident trends and reduce toil through automation.
  • Strategic Automation and AI: Apply automation and AI to improve triage, runbook execution, and anomaly detection.
  • Observability, Monitoring, and Alert Quality: Improve instrumentation, SLIs, dashboards, and actionable alerts with better thresholds and aggregation.
  • SLOs, Error Budgets, and Reliability Governance: Define/ steward SLOs/SLIs and manage error budgets.
  • Capacity Engineering and Scalability: Drive capacity testing, scaling actions, and headroom reporting.
  • Change Quality and ORR Gatekeeping: Oversee change quality and validate preparedness before go-live.
  • Documentation, Runbooks, and Training: Improve docs and train developers on SRE practices.
  • KPIs and Reporting: Publish KPI/OKR dashboards and executive-ready reports.

Skills

Incident management
SRE fundamentals
Cross-region coordination
Automation & AI
Stakeholder communications

Tools

Prometheus
Grafana
OpenTelemetry
ELK
PagerDuty
Opsgenie
Terraform
CloudFormation
Ansible

Job description

Goldman Sachs is seeking a Senior Site Reliability Engineer (SRE) specializing in incident management, escalation discipline, and automation. You will partner with front-office desks, engineering, and ABO to improve desk readiness and reduce toil through automation and AI.

You will lead cross-region handoffs, governance of error budgets, and establish clear SRE KPIs. The role emphasizes observability, scalable systems, and reliable incident communications across globally distributed

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Global Incident Commander & Automation
Senior SRE: Global Incident Commander & Automation

Goldman Sachs • Singapore

On-site
SGD 150,000 - 210,000
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore

Goldman Sachs • Singapore

On-site
SGD 150,000 - 210,000
Senior Site Reliability Engineer: AI-Driven Platform Automation
Senior Site Reliability Engineer: AI-Driven Platform Automation

SGX Group • Singapore

On-site
SGD 90,000 - 130,000
Global Banking & Markets SRE – AI-Driven, Cloud-Native
Global Banking & Markets SRE – AI-Driven, Cloud-Native

goldman sachs services (singapore) pte. ltd. • Singapore

On-site
SGD 180,000 - 300,000
Platform SRE Lead — Reliability, AI-Driven Incident Response
Platform SRE Lead — Reliability, AI-Driven Incident Response

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Senior SRE Lead: Resiliency, Incidents & Observability
Senior SRE Lead: Resiliency, Incidents & Observability

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 180,000
Global Banking & Markets, Site Reliability Engineer, Executive Director, Singapore
Global Banking & Markets, Site Reliability Engineer, Executive Director, Singapore

goldman sachs services (singapore) pte. ltd. • Singapore

On-site
SGD 180,000 - 300,000
Chief SRE & Reliability Leader, 24/7 Ops Governance
Chief SRE & Reliability Leader, 24/7 Ops Governance

DBS Bank • Singapore

On-site
SGD 300,000 - 520,000
Global SRE & Reliability Transformation Leader
Global SRE & Reliability Transformation Leader

DBS BANK LTD. • Singapore

On-site
SGD 300,000 - 420,000
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore Singapore · Singapore · Vice President
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore Singapore · Singapore · Vice President

Goldman Sachs Bank AG • Singapore

On-site
SGD 120,000 - 180,000