Engineering Division - SRE Platforms - Software Engineering - Vice President

Goldman Sachs

Hyderabad

On-site

INR 3,500,000 - 5,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Goldman Sachs is seeking an experienced Site Reliability Engineer (Vice President) to champion the availability, reliability, and scalability of our critical platform services. You will combine software and systems expertise to architect and operate large-scale, fault-tolerant systems across on-prem and cloud environments.

You will mentor senior engineers, drive continuous improvement, and promote advanced SRE practices across the organization while collaborating with executive stakeholders and

Qualifications

  • Minimum 6+ years of hands-on experience in Site Reliability Engineering.

Responsibilities

  • Drive strategic availability, scalability, and performance of mission-critical applications.
  • Lead design and implementation of resilient infrastructure and architectures.
  • Develop tooling and automation to eliminate toil and improve deployment reliability.
  • Lead incident response and post-mortem analyses with long-term preventive measures.
  • Collaborate with development teams to bake reliability into design from inception.
  • Define observability strategies including monitoring, logging, tracing, and SLIs/SLOs.
  • Mentor engineers and lead technical initiatives and SDLC best practices.

Skills

SRE
Cloud
Automation

Tools

Monitoring
Logging

Job description

Description Site Reliability Engineer - Vice President

Site Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run scalable, massively distributed, fault-tolerant systems.

At Goldman Sachs, SRE is responsible for improving the availability and reliability of the firms most critical platform services and ensures they meet the requirements of our internal and external users. It is also responsible for the firmwide policies and standards focused on firms digital resilience. We are looking for engineers who are motivated to collaborate with our businesses to build and run sustainable production systems, which can evolve and adapt to changes in our fast-paced, global business environment.

The SRE team develops and maintains platforms and tools which help other Engineering teams in Goldman Sachs to build and operate reliable and resilient systems. These systems span on-premises datacenters and multiple public cloud environments. The platforms we offer include central logging, monitoring, agents and alerting and we provide tools to drive adoption and improvements to capacity planning, operational readiness assessments, production incident postmortems, SLIs / SLOs, and deployment automation including canary releases.

The products and services we provide to our internal customers are used by thousands of engineers every day. We believe that reliability is the most important feature of any system, and we are devoted to giving our engineers the platforms and tools they need to build and operate reliable products.

Role Overview As a Site Reliability Engineer (SRE) at Goldman Sachs, you will be a pivotal leader in ensuring the availability, reliability, and scalability of the firm's most critical platform applications and services. You will combine deep software and systems engineering expertise to architect, build, and run large-scale, massively distributed, fault-tolerant systems.

This role involves providing technical leadership, mentoring senior engineers, and collaborating closely with internal teams and executive stakeholders to build and operate sustainable production systems that can adapt to our dynamic global business environment. You will drive a culture of continuous improvement, championing the adoption of advanced SRE principles and best practices across the organization.

Responsibilities
  • Strategic Reliability & Performance: Drive the strategic direction for availability, scalability, and performance of mission-critical applications and platform services, ensuring alignment with firm-wide objectives.
  • Architectural Leadership: Lead the design, build, and implementation of highly available, resilient, and scalable infrastructure and application architectures.
  • Advanced Automation & Tooling: Architect and develop sophisticated platforms, tools, and automation solutions to eliminate toil, optimize operational workflows, and enhance deployment processes across the enterprise.
  • Complex Incident Management & Post-Mortem Analysis: Lead critical incident response, conduct in-depth root cause analysis for systemic issues, and implement long-term preventative measures to significantly enhance system stability and resilience.
  • System Design & Capacity Planning: Partner with development teams to embed reliability into application design from inception, provide expert system design consulting, and lead comprehensive capacity planning initiatives for future growth.
  • Observability & Insights: Define and implement advanced monitoring, high volume logging with multi-user query capabilities, and tracing strategies to provide deep, actionable insights into application performance, infrastructure health, and user experience.
  • Technical Vision & Mentorship: Provide technical vision, lead complex technical projects, conduct rigorous code reviews, enforce SDLC best practices, and actively mentor and develop senior and staff-level engineers.
  • Technology Evaluation & Adoption: Stay at the forefront of industry trends and advancements, evaluating and integrating cutting-edge tools and frameworks to significantly improve operational efficiency and reliability.
  • On-Call Leadership: Participate in and lead on-call rotations, providing expert guidance and hands‑on support for critical system incidents.
Qualifications

Experience: Minimum of 6+ years of hands‑on experience in Site Reliability Engineering, with a proven track record in architecting, designing, building, and maintaining highly available, scalable, and fault-tole.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Vice President - Site Reliability Engineering (SRE)
Vice President - Site Reliability Engineering (SRE)

Goldman Sachs Services Pvt Ltd • Bengaluru

On-site
INR 3,000,000 - 7,000,000
Site Reliability Engineer - Vice President
Site Reliability Engineer - Vice President

Citi • Pune District

On-site
INR 4,000,000 - 5,500,000
The Core Engineering-L2-Bengaluru-Associate-Software Engineering
The Core Engineering-L2-Bengaluru-Associate-Software Engineering

Goldman Sachs • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer-Vice President
Site Reliability Engineer-Vice President

Citi Bank • Pune District

On-site
INR 1,800,000 - 3,200,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
VP_SRE (Only from Investment Bank or Product Company)
VP_SRE (Only from Investment Bank or Product Company)

Mancer Consulting Services • Mumbai Suburban

On-site
INR 2,000,000 - 3,000,000
The Core Engineering-L2-Bengaluru-Vice President-Software Engineering Bengaluru · India · Vice President
The Core Engineering-L2-Bengaluru-Vice President-Software Engineering Bengaluru · India · Vice President

Goldman Sachs Bank AG • Bengaluru

On-site
INR 1,500,000 - 2,300,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000