SRE Platforms Engineer: Reliability, Observability & Scale

The Goldman Sachs Group

Dallas (TX)

On-site

USD 110,000 - 140,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

The Goldman Sachs Group is seeking a Site Reliability Engineer to help design, build, and operate large-scale, fault-tolerant services. You will collaborate with software engineers to ensure uptime and performance of critical platforms used across the firm.

In this role, you will implement SLOs, monitor system health, participate in incident response, and drive automation and reliability improvements across cloud and on‑prem environments.

Qualifications

  • Minimum of 2+ years of hands-on SRE experience building highly available, scalable, fault-tolerant systems.
  • BS degree in Computer Science or related technical field.
  • Proficiency in Go, Python, C, C++, Java, Perl, Ruby or shell scripting.
  • Experience with UNIX internals and/or networking.

Responsibilities

  • Balance feature velocity and reliability with well-defined SLOs.
  • Run the Production environment by monitoring availability and system health.
  • Drive incident management and blameless post-mortems culture.
  • Partner with development teams to improve services via testing and release procedures.
  • Participate in system design, platform management, and capacity planning.
  • Create sustainable systems and services through automation.
  • Champion reliability and resilience engineering practices across the firm.

Skills

2+ years SRE experience
Go
Python
C
C++
Java
Perl
Ruby
Shell scripting

Education

BS in Computer Science or related field

Job description

The Goldman Sachs Group is seeking a Site Reliability Engineer to help design, build, and operate large-scale, fault-tolerant services. You will collaborate with software engineers to ensure uptime and performance of critical platforms used across the firm.

In this role, you will implement SLOs, monitor system health, participate in incident response, and drive automation and reliability improvements across cloud and on‑prem environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Platform Engineer: Reliability & Observability
SRE Platform Engineer: Reliability & Observability

Goldman Sachs • Dallas (WV)

On-site
USD 120,000 - 180,000
Platform SRE Engineer — Reliability & Observability
Platform SRE Engineer — Reliability & Observability

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
None
VP, SRE Platforms — Scale, Reliability & Automation
VP, SRE Platforms — Scale, Reliability & Automation

The Goldman Sachs Group • Dallas (TX)

On-site
USD 180,000 - 280,000
Senior SRE - Global Markets & AI Ops
Senior SRE - Global Markets & AI Ops

Socket.dev • New York (NY)

On-site
USD 150,000 - 300,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
None
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Goldman Sachs • Dallas (WV)

On-site
USD 120,000 - 180,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 110,000 - 140,000
VP Site Reliability Engineer — Global Markets
VP Site Reliability Engineer — Global Markets

Goldman Sachs • New York (NY)

On-site
USD 150,000 - 250,000
VP, SRE Platforms & Reliability
VP, SRE Platforms & Reliability

Goldman Sachs, Inc. • Dallas (TX)

On-site
USD 210,000 - 260,000
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 180,000 - 280,000