Site Reliability Engineering (SRE), The Core Engineering, Analyst, Dallas

Goldman Sachs Group, Inc.

Dallas (TX)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Goldman Sachs is seeking an experienced Site Reliability Engineer in Dallas to help build and operate highly reliable platforms. This VP-level role focuses on SRE principles, scalable systems, and production excellence across a large engineering organization.

The role emphasizes defining SLOs/SLIs, reducing toil through automation, and driving blameless post-mortems to prevent incidents, while collaborating with product teams on resilient, observable infrastructure.

Qualifications

  • Proficiency in at least one major language (Java/Python/Node.js) for tooling and automation.
  • Hands-on experience with IaC (Terraform/Ansible/CloudFormation).
  • Deep understanding of containers and Kubernetes; incl. service meshes.
  • Experience with AWS/GCP/Azure; resilient cloud-native architectures.
  • Observability stack: Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, CloudWatch.
  • Linux environment development, algorithms, data structures, and software design.
  • Networking and load balancing in distributed systems.

Responsibilities

  • Establish SLOs, SLIs, and error budgets.
  • Collaborate to architect highly available, fault-tolerant systems.
  • Reduce operational toil with automation and self-service tooling.
  • Improve production readiness via load testing, capacity forecasting, reliability reviews.
  • Lead complex multi-system production incident responses with blameless post-mortems.
  • Design healthy on-call models, clear escalation paths, and balanced pager responsibilities.

Skills

Java/Python/Node.js
Terraform/Ansible/CloudFormation
Docker/Kubernetes
Cloud: AWS/GCP/Azure
Observability stack
Linux/Algorithms
Networking/Load balancing

Education

Bachelor’s degree in Computer Science or related field

Tools

Terraform
Docker
Kubernetes
OpenTelemetry
ELK/Datadog/Prometheus

Job description

Site Reliability Engineering (SRE), The Core Engineering, Analyst, Dallas
Job Description

WHAT WE DO

Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help engineer highly reliable, observable, and resilient platforms that support critical business services at scale. You will collaborate with multiple engineering teams to continually improve our production system architecture, facilitate fast delivery of new services, and reduce downtime.

This role is for software engineers who enjoy solving complex distributed system problems, building tools and platforms that make teams more effective, and championing SRE principles (such as SLOs, error budgets, and blameless post-mortems) across a large engineering organization.

Key Responsibilities
  • Partner with engineering leadership to establish service level objectives (SLOs), service level indicators (SLIs), and error budgets.
  • Collaborate with product developers to architect highly available, fault-tolerant, and self-healing systems. Conduct architectural reviews and introduce patterns like circuit breakers, graceful degradation, and rate limiting.
  • Reduce operational toil by building automation, tooling, and self-service capabilities that remove repetitive manual work.
  • Improve production readiness through load testing, performance tuning, capacity forecasting, and reliability reviews.
  • Lead the response to complex, multi-system production incidents. Facilitate blameless post-mortems to identify root causes and drive long-term preventative actions.
  • Promote sustainable operations by helping design healthy on‑call models, clear escalation paths, and balanced pager responsibilities.
WHAT WE ARE LOOKING FOR
Core Technical Skills
  • Strong proficiency in at least one major programming language (e.g., Java, Python, or Node.js) with a focus on writing clean, maintainable code for tooling and automation.
  • Hands‑on experience with Infrastructure as Code (IaC) frameworks such as Terraform, Ansible, or CloudFormation.
  • Deep understanding of containerization and orchestration technologies, specifically Docker and Kubernetes (K8s), including service meshes and ingress controllers.
  • Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud‑native architectures.
  • Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch)
  • Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software design.
  • Knowledge of networking protocols and load balancing strategies in a distributed systems environment.
Core Competencies & Soft Skills
  • Ability to analyze complex, distributed systems holistically and understand how individual components interact under load.
  • Strong interpersonal skills to collaborate with product developers, influence architectural decisions,prioritize toil reduction, and drive SRE adoption without direct authority.
  • Ability to translate complex technical issues into clear, actionable insights for both technical and non-technical stakeholders.
  • Highly motivated, pro‑active and capable of multi‑tasking under pressure in a fast‑paced environment without compromising quality.
  • Commitment to fostering a blameless culture where failures are treated as opportunities to learn and improve systems.
  • Interest in financial markets and technology.
Preferred Qualifications
  • Bachelor’s degree in Computer Science, System Engineering, or a related technical field that involves programming.
  • 7 to 10 years of experience
ABOUT GOLDMAN SACHS

The Goldman Sachs Group, Inc. is a leading global investment banking, securities and investment management firm that provides a wide range of financial services to a substantial and diversified client base that includes corporations, financial institutions, governments and individuals. Founded in 1869, the firm is headquartered in New York and maintains offices in all major financial centers around the world.

Healthcare & Medical Services

We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has a number of opportunities to grow professionally and personally

We offer competitive vacation policies based on employee level and office location. We promote time off from work to recharge by providing generous vacation entitlements and a minimum of three weeks expected vacation usage each year.

Financial Wellness & Retirement

We assist employees in saving and planning for retirement, offer financial support for higher education, and provide a number of benefits to help employees prepare for the unexpected. We offer live financial education and content on a variety of topics to address the spectrum of employees’ priorities.

Health

We offer a medical advocacy service for employees and family members facing critical health situations, and counseling and referral services through the Employee Assistance Program (EAP). We provide Global Medical, Security and Travel Assistance and a Workplace Ergonomics Program. We also offer state‑of-the‑art on‑site health centers in certain offices.

Fitness

To encourage employees to live a healthy and active lifestyle, some of our offices feature on‑site fitness centers. For eligible employees we typically reimburse fees paid for a fitness club membership or activity (up to a pre‑approved amount).

We offer on‑site child care centers that provide full‑time and emergency back‑up care, as well as mother and baby rooms and homework rooms. In every office, we provide advice and counseling services, expectant parent resources and transitional programs for parents returning from parental leave. Adoption, surrogacy, egg donation and egg retrieval stipends are also available.

Benefits at Goldman Sachs

Read more about the full suite of class‑leading benefits our firm has to offer.

Learn More

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Vice President - Site Reliability Engineering (SRE) – The Core Engineering New York · · Vice President
Vice President - Site Reliability Engineering (SRE) – The Core Engineering New York · · Vice President

Goldman Sachs Bank AG • Northern (KY), New York (NY)

Hybrid
USD 180,000 - 240,000
Vice President - Site Reliability Engineering (SRE) – The Core Engineering
Vice President - Site Reliability Engineering (SRE) – The Core Engineering

Goldman Sachs Group, Inc. • New York (NY)

On-site
USD 190,000 - 270,000
Site Reliability Engineer, Global Banking & Markets, Vice President
Site Reliability Engineer, Global Banking & Markets, Vice President

Goldman Sachs Group, Inc. • Northern (KY), New York (NY)

Hybrid
USD 150,000 - 250,000
Site Reliability Engineer, Global Banking & Markets, Frontline Production Engineering New York [...]
Site Reliability Engineer, Global Banking & Markets, Frontline Production Engineering New York [...]

Goldman Sachs Bank AG • New York (NY)

On-site
USD 130,000 - 250,000
Site Reliability Engineer, Global Banking & Markets, Vice President
Site Reliability Engineer, Global Banking & Markets, Vice President

Goldman Sachs Bank AG • New York (NY)

On-site
USD 150,000 - 250,000
Software Engineering - SRE Data - Software Engineer - Associate - Dallas
Software Engineering - SRE Data - Software Engineer - Associate - Dallas

Goldman Sachs Group, Inc. • Dallas (TX)

On-site
USD 120,000 - 180,000
Vice President - Site Reliability Engineering (SRE) - The Core Engineering
Vice President - Site Reliability Engineering (SRE) - The Core Engineering

The Goldman Sachs Group • New York (NY)

On-site
USD 250,000 - 320,000
Site Reliability Engineer, Global Banking & Markets, Frontline Production Engineering
Site Reliability Engineer, Global Banking & Markets, Frontline Production Engineering

The Goldman Sachs Group • New York (NY)

On-site
USD 130,000 - 250,000
Software Engineer, Global Banking & Markets, Trading Technology Salt Lake City · · Associate
Software Engineer, Global Banking & Markets, Trading Technology Salt Lake City · · Associate

Goldman Sachs Bank AG • Salt Lake City (UT)

On-site
USD 110,000 - 165,000
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 180,000 - 280,000