Global Banking & Markets, Site Reliability Engineer, Vice President, Singapore
Job Description
You will join theSite Reliability Engineering (SRE)team supporting Global Banking and Markets, ensuring the availability, resilience, and performance of core business services that underpin a global 24×7 trading operation. Working across Global Markets front, middle, and back office functions, you will engineer reliability into every layer of our platform — balancing stringent non-functional demands foravailability, latency, and resiliencewhile meeting complex, evolving business requirements.
Critically, you will operate at the intersection ofreliability engineering, cloud infrastructure, and AI-driven operations. Using Goldman Sachs' AI tooling and agentic assistants, you will accelerate incident diagnosis, automate operational toil, comprehend large legacy codebases, and raise the bar for production-quality automation across the software and reliability lifecycle. Above all, you will bringstrong risk acumenand the ability toconnect the right people across the organizationto resolve problems quickly and decisively.
Your Impact
- Own reliability outcomes:Define and defend Service Level Objectives (SLOs), error budgets, and reliability standards for critical trading services, with risk always front of mind.
- Reduce risk and toil:Identify systemic risks before they materialize, automate away repetitive operational work, and strengthen the resilience posture of the platform.
- Connect and communicate:Act as a trusted coordinator during incidents — rapidly mobilizing the right engineers, domain experts, and stakeholders across a globally distributed organization, and communicating clearly with both technical and non-technical audiences.
- Multiply your output with AI:Orchestrate AI coding and operations agents to accelerate root-cause analysis, remediation, and automation while maintaining mastery, quality, and production fitness over all AI-generated work.
- Build for the future:Design and operate high-availability, multi-region, event-driven services on a modern cloud-native platform, setting the reliability and architectural standard for years to come.
What You Will Do
- Design, build, and operatehigh-availability, multi-region, cloud-native serviceswith security and comprehensive observability (metrics, distributed tracing, structured logging) built in at every layer.
- Establish and manageSLIs, SLOs, and error budgets; drive blameless post-incident reviews and translate findings into durable engineering improvements.
- Lead incident response for latency-sensitive, high-throughput trade lifecycle systems — quickly diagnosing issues, coordinating cross-functional responders, and communicating status to stakeholders.
- Developevent-driven architectures, multi-stage processing pipelines, and optimized data pathsfor high-throughput trade lifecycle management.
- Applystrong risk acumento change management, capacity planning, and resilience testing (chaos engineering, failover, and BCP drills).
- Partner with engineers, domain experts, and global stakeholders to understand production processes, challenge entrenched assumptions in a cloud-centric, AI-driven world, and drive modernization.
- Multiply your impact with a modern,AI-centric toolchain, orchestrating AI agents across the SDLC and operations to rapidly comprehend large codebases, generate production-quality automation, and accelerate delivery.
Required Qualifications
- Minimum 8 yearsof professional software / reliability engineering experience, with strong command of at least one major language (Java 17+ preferred), including concurrency, collections, and modern language features.
- Demonstrated risk acumen— the ability to identify, quantify, and mitigate operational and technical risk in a regulated financial services environment.
- Excellent communication and stakeholder-coordination skills— proven ability to connect the right people quickly and drive resolution across geographically distributed, technical and non-technical audiences.
- Proven experience runninghigh-availability production environments: SLIs/SLOs, error budgets, on-call, incident command, and post-incident reviews.
- Strong understanding ofcloud infrastructure(GCP, AWS), container orchestration (Kubernetes, Docker), and infrastructure-as-code.
- Working knowledge ofAI models and AI-assisted engineering tools(e.g., Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist), including the ability to govern AI agents, critically assess their output, and maintain quality over AI-generated work.
- Experience buildingevent-driven and distributed systems, including messaging platforms (e.g., Apache Kafka), delivery guarantees, and resilience strategies.
- StrongSDLC and automation practices: version control, CI/CD pipelines, automated build/test/deploy workflows, and code quality tooling.
- Solidobservability discipline: application instrumentation, distributed tracing, structured logging, and metrics-driven operations.
- Ability to rapidly navigate, understand, and debug large and unfamiliar codebases — with and without AI assistance.
Preferred Qualifications
Experience with a meaningful subset of the following is highly valued:
- Reliability & Operations:Chaos engineering, capacity planning, load/performance testing, and production support in high-availability, latency-sensitive environments.
- Frameworks & Architecture:Spring Boot, gRPC / Protocol Buffers, integration/orchestration frameworks (e.g., Apache Camel, Spring Integration), and pipeline/adapter patterns (retry, dead-letter queues, error isolation).
- Cloud & Infrastructure:Cloud platforms (GCP, AWS), Kubernetes/Docker, JVM tuning for containerized workloads, and infrastructure-as-code (Terraform, Helm).
- AI & Automation:Applying AI models to operational use cases — anomaly detection, log analysis, automated remediation, and agentic operations.
- Observability & Operations:Prometheus, Grafana, OpenTelemetry, and SLO tooling.
- Data & Performance:Data modeling, SQL/NoSQL databases, caching strategies, and performance optimization in latency-sensitive systems.
- Security:Enterprise security patterns; authentication protocols, mutual TLS, secrets management, and certificate rotation.
- Domain Knowledge:Equities, post-trade, or financial services experience; trade lifecycle concepts, position management, reconciliation, and multi-system migration environments.
- Other:Asynchronous / non-blocking I/O frameworks (e.g., Vert.x, Netty), multi-region / BCP architectures, and open-source contribution experience.
ABOUT GOLDMAN SACHS
At Goldman Sachs, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world.
We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has a number of opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. Learn more about our culture, benefits, and people at GS.com/careers.
We’re committed to finding reasonable accommodations for candidates with special needs or disabilities during our recruiting process. Learn more: https://www.goldmansachs.com/careers/footer/disability-statement.html
Goldman Sachs is an equal opportunity employer and does not discriminate on the basis of race, color, religion, sex, national origin, age, veterans status, disability, or any other characteristic protected by applicable law.
Job Info
- Job Identification 180983
- Job Category Vice President
Healthcare & Medical Services
We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has a number of opportunities to grow professionally and personally.
We offer competitive vacation policies based on employee level and office location. We promote time off from work to recharge by providing generous vacation entitlements and a minimum of three weeks expected vacation usage each year.
Financial Wellness & Retirement
We assist employees in saving and planning for retirement, offer financial support for higher education, and provide a number of benefits to help employees prepare for the unexpected. We offer live financial education and content on a variety of topics to address the spectrum of employees’ priorities.
Health
We offer a medical advocacy service for employees and family members facing critical health situations, and counseling and referral services through the Employee Assistance Program (EAP). We provide Global Medical, Security and Travel Assistance and a Workplace Ergonomics Program. We also offer state-of-the-art on-site health centers in certain offices.
Fitness
To encourage employees to live a healthy and active lifestyle, some of our offices feature on-site fitness centers. For eligible employees we typically reimburse fees paid for a fitness club membership or activity (up to a pre-approved amount).
We offer on-site child care centers that provide full-time and emergency back-up care, as well as mother and baby rooms and homework rooms. In every office, we provide advice and counseling services, expectant parent resources and transitional programs for parents returning from parental leave. Adoption, surrogacy, egg donation and egg retrieval stipends are also available.
Benefits at Goldman Sachs
Read more about the full suite of class-leading benefits our firm has to offer.
Learn More