Get more replies from employers
Send a job-specific resume in minutes.
Goldman Sachs is seeking a Site Reliability Engineer with 2+ years of hands-on experience to build and operate highly available, scalable, and fault-tolerant systems. This role focuses on reliability, monitoring, incident management, and collaboration with cross-functional teams to improve services through rigorous testing and robust release procedures.
The ideal candidate has strong coding skills in Go, Python, C/C++, Java, and shell scripting, plus experience with UNIX systems and networking.
Site Reliability Engineering (SRE) is an engineering discipline that combines software development and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. At Goldman Sachs, SRE is responsible for the availability and reliability of our firm's most critical platform services and ensures they meet the requirements of our internal and external users. We also develop and operate the observability platforms that all other engineering teams use to make their services reliable. We look for engineers who are motivated to collaborate with other engineering teams and our businesses to build and run sustainable production systems, which can evolve and adapt to changes in our fast-paced, global business and regulatory environment.