Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas

The Goldman Sachs Group

Dallas (TX)

On-site

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Goldman Sachs Group is seeking an experienced Site Reliability Engineer (Vice President) to lead the availability, reliability, and scalability of its critical platform services.

You will architect, build, and operate massively distributed systems across on‑prem and cloud environments, mentoring engineers and partnering with executives to drive best‑practice reliability and efficiency.

Qualifications

  • 6+ years of hands‑on SRE experience in enterprise environments.
  • Strong programming skills in Java, Python, or Go.
  • Extensive cloud experience (AWS, GCP) and container orchestration (Docker, Kubernetes).
  • IaC with Terraform/CloudFormation and configuration management tools.

Responsibilities

  • Drive strategic availability, scalability, and performance of mission‑critical services.
  • Lead architectural design of highly available infrastructure and applications.
  • Build automation platforms to reduce toil and improve deployment workflows.
  • Own incident management with deep root‑cause analysis and preventative measures.
  • Collaborate with development teams to embed reliability in design and capacity planning.
  • Define observability strategies with monitoring, logging, and tracing.

Skills

Java
Python
Go
Cloud AWS
Kubernetes
Docker
IaC Terraform
CI/CD
Linux
Distributed systems

Education

Advanced degree in CS

Tools

Terraform
CloudFormation
Jenkins
GitLab
Prometheus
Grafana
ELK stack
Datadog

Job description

Site Reliability Engineer - Vice President

Site Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run scalable, massively distributed, fault‑tolerant systems. At Goldman Sachs, SRE is responsible for improving the availability, reliability, and scalability of the firm's most critical platform services and for making sure they meet the requirements of internal and external users. It also establishes firm‑wide policies and standards focused on digital resilience. We are looking for engineers who are motivated to collaborate with businesses to build and run sustainable production systems that can evolve and adapt to changes in our fast‑paced, global environment.

The SRE team develops and maintains platforms and tools that help other engineering teams at Goldman Sachs to build and operate reliable and resilient systems. These systems span on‑premises datacenters and multiple public cloud environments. The platforms we offer include central logging, monitoring, agents, and alerting, and we provide tools to drive adoption and improvements in capacity planning, operational readiness assessments, production incident post‑mortem analysis, SLIs/SLOs, and deployment automation including canary releases.

The products and services we provide to our internal customers are used by thousands of engineers every day. We believe that reliability is the most important feature of any system, and we are devoted to giving our engineers the platforms and tools they need to build and operate reliable products.

Role Overview

As a Site Reliability Engineer (SRE) at Goldman Sachs, you will be a pivotal leader in ensuring the availability, reliability, and scalability of the firm's most critical platform applications and services. You will combine deep software and systems engineering expertise to architect, build, and run large‑scale, massively distributed, fault‑tolerant systems. This role involves providing technical leadership, mentoring senior engineers, and collaborating closely with internal teams and executive stakeholders to build and operate sustainable production systems that can adapt to our dynamic global business environment. You will drive a culture of continuous improvement, championing the adoption of advanced SRE principles and best practices across the organization.

Responsibilities
  • Strategic Reliability & Performance: Drive the strategic direction for availability, scalability, and performance of mission‑critical applications and platform services, ensuring alignment with firm‑wide objectives.
  • Architectural Leadership: Lead the design, build, and implementation of highly available, resilient, and scalable infrastructure and application architectures.
  • Advanced Automation & Tooling: Architect and develop sophisticated platforms, tools, and automation solutions to eliminate toil, optimize operational workflows, and enhance deployment processes across the enterprise.
  • Complex Incident Management & Post‑Mortem Analysis: Lead critical incident response, conduct in‑depth root‑cause analysis for systemic issues, and implement long‑term preventative measures to significantly enhance system stability and resilience.
  • System Design & Capacity Planning: Partner with development teams to embed reliability into application design from inception, provide expert system design consulting, and lead comprehensive capacity‑planning initiatives for future growth.
  • Observability & Insights: Define and implement advanced monitoring, high‑volume logging with multi‑user query capabilities, and tracing strategies to provide deep, actionable insights into application performance, infrastructure health, and user experience.
  • Technical Vision & Mentorship: Provide technical vision, lead complex technical projects, conduct rigorous code reviews, enforce SDLC best practices, and actively mentor and develop senior and staff‑level engineers.
  • Technology Evaluation & Adoption: Stay at the forefront of industry trends and advancements, evaluating and integrating cutting‑edge tools and frameworks to significantly improve operational efficiency and reliability.
  • On‑Call Leadership: Participate in and lead on‑call rotations, providing expert guidance and hands‑on support for critical system incidents.
Qualifications
  • Experience: Minimum of 6+ years of hands‑on experience in Site Reliability Engineering, with a proven track record in architecting, designing, building, and maintaining highly available, scalable, and fault‑tolerant systems at an enterprise level.
  • Technical Proficiency:
    • Exceptional programming skills in one or more major languages such as Java, Python, Go with a focus on building robust, scalable software.
    • Extensive hands‑on experience with cloud platforms (e.g., AWS, GCP) and deep expertise in containerization and orchestration technologies (e.g., Docker, Kubernetes).
    • Mastery of Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation) and configuration management tools (e.g., Puppet, Chef, Ansible).
    • Advanced proficiency in Prompt Engineering and Retrieval‑Augmented Generation (RAG) architectures to automate complex SRE workflows, such as the generation of Infrastructure as Code (IaC), dynamic runbooks, and incident response summaries.
    • Profound understanding of Linux internals, networking, distributed systems, and advanced system performance tuning.
    • Expertise in designing and implementing comprehensive monitoring, alerting, logging, and tracing solutions (e.g., Prometheus, Grafana, ELK stack, Datadog, PagerDuty).
    • Deep experience with CI/CD tools and practices (e.g., Jenkins, GitLab, Maven).
    • Strong foundation in databases and distributed systems.
    • Exceptional problem‑solving abilities and analytical skills, with a track record of resolving complex technical challenges.
  • Preferred Experience:
    • Experience with distributed databases such as ElasticSearch.
    • Experience with working on GCP BigQuery.
    • Experience with messaging systems such as Kafka.
  • Education: Advanced degree (Bachelor's, Master's, or PhD) in Computer Science or a related technical field involving coding and/or systems engineering, or equivalent practical experience.
  • Soft Skills: Superior communication, collaboration, and interpersonal skills, with the ability to influence technical direction, lead cross‑functional initiatives, and effectively engage with global teams and executive leadership. Proven ability to work independently, manage multiple complex stakeholders, and drive significant organizational change.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 110,000 - 140,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
None
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Goldman Sachs • Dallas (WV)

On-site
USD 120,000 - 180,000
Engineering – SRE Platforms – Software Engineer – Vice President – Dallas | Dallas, TX, USA
Engineering – SRE Platforms – Software Engineer – Vice President – Dallas | Dallas, TX, USA

Goldman Sachs, Inc. • Dallas (TX)

On-site
USD 210,000 - 260,000
Site Reliability Engineering (SRE), The Core Engineering, Vice President, Dallas
Site Reliability Engineering (SRE), The Core Engineering, Vice President, Dallas

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
Site Reliability Engineer, Global Banking & Markets, Vice President
Site Reliability Engineer, Global Banking & Markets, Vice President

Socket.dev • New York (NY)

On-site
USD 150,000 - 300,000
Site Reliability Engineer, Global Banking & Markets, Vice President
Site Reliability Engineer, Global Banking & Markets, Vice President

Goldman Sachs • New York (NY)

On-site
USD 150,000 - 250,000
Senior SRE, Compliance Engineering & DevOps
Senior SRE, Compliance Engineering & DevOps

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
Site Reliability Engineer, Global Banking & Markets, Frontline Production Engineering
Site Reliability Engineer, Global Banking & Markets, Frontline Production Engineering

The Goldman Sachs Group • New York (NY)

On-site
USD 130,000 - 250,000
Site Reliability Engineer, Global Banking & Markets, Vice President
Site Reliability Engineer, Global Banking & Markets, Vice President

The Goldman Sachs Group • New York (NY)

On-site
USD 150,000 - 250,000