Get more replies from employers
Send a job-specific resume in minutes.
The Goldman Sachs Group is seeking a Site Reliability Engineer to help design, build, and operate large-scale, fault-tolerant services. You will collaborate with software engineers to ensure uptime and performance of critical platforms used across the firm.
In this role, you will implement SLOs, monitor system health, participate in incident response, and drive automation and reliability improvements across cloud and on‑prem environments.
Site Reliability Engineering (SRE) is an engineering discipline that combines software development and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. At Goldman Sachs, SRE is responsible for the availability and reliability of our firm's most critical platform services and ensures they meet the requirements of our internal and external users. We also develop and operate the observability platforms that all other engineering teams use to make their services reliable. We look for engineers who are motivated to collaborate with other engineering teams and our businesses to build and run sustainable production systems, which can evolve and adapt to changes in our fast-paced, global business and regulatory environment.