Stand out for this role — generate a tailored resume and cover letter in about a minute.
DBS Bank Ltd in Singapore seeks a Senior VP, Site Reliability Engineering to lead and oversee a 24/7 SRE function across critical platforms including OpenShift, Kubernetes, VMware, Windows, Linux, databases and mainframe.
You will drive resilience, automation, and operational excellence across hybrid cloud and on-premises environments, while aligning with risk and regulatory expectations and partnering with security teams to harden infrastructure.
The SVP, Site Reliability Engineering (SRE), will lead and oversee the 24/7 infrastructure operations and reliability engineering function across critical platforms including Hypervisors (VPC, EPC, OPC), OpenShift , Windows, Databases, TWS, and Mainframe environments. This role is responsible for driving resilience, scalability, automation, and operational excellence across hybrid cloud and on-premises environments, while ensuring alignment with business, risk, and regulatory expectations.
Lead and manage a distributed 24/7 SRE infrastructure team, including shift-based operations and command center functions Define and execute the SRE strategy aligned to enterprise technology and business priorities Establish strong governance across incident, problem, change, release, and capacity management Drive SLA/SLO/SLI frameworks to ensure service reliability and performance targets
Oversee end-to-end reliability of infrastructure platforms: Cloud & Container: VPC, OpenShift, Kubernetes Compute & Virtualization: Hypervisors (VMware/others), private cloud platforms Enterprise Platforms: Windows, Unix/Linux, TWS, Mainframe, Databases Ensure high availability, resilience, and disaster recovery readiness across all critical systems Own infrastructure lifecycle including capacity planning, patching, upgrades, and decommissioning
Champion SRE principles including error budgets, toil reduction, and automation-first mindset Drive end-to-end observability strategy (monitoring, logging, tracing) Lead initiatives to reduce MTTR, incident volume, and manual operational effort Scale automation across deployment, patching, incident resolution, and self-healing capabilities
Ensure 24/7 monitoring, incident response, and recovery processes are robust and continuously improved Lead major incident management and command bridge coordination for critical outages Conduct RCA, trend analysis, and preventive engineering improvements Embed ITIL best practices across service management processes
Identify infrastructure risks and drive proactive mitigation strategies Ensure compliance with regulatory, audit, and internal security requirements Partner with security teams on hardening, vulnerability management, and access controls
Collaborate with application, DevOps, security, architecture, and business teams to improve system reliability Provide leadership in large-scale transformation programs (cloud adoption, infra modernization, SRE maturity) Act as a key interface with senior management and external stakeholders
Build and develop a high-performing SRE organization across L1/L2/L3 layers Drive fungibility, cross-skilling, and leadership development within the team Mentor senior leaders and establish clear career progression frameworks
18+ years of experience in IT infrastructure, SRE, or production operations Proven leadership in managing large-scale 24/7 infrastructure teams in banking/financial services Strong experience in hybrid cloud, data center, and enterprise platforms
Deep expertise in: Cloud platforms (private/public cloud architectures) Container platforms (OpenShift/Kubernetes) Hypervisors & virtualization technologies Operating systems (Windows, Linux/Unix) Databases (MariaDB, Postgres, MSSQL, Redis, DB2) Enterprise scheduling & legacy systems (TWS, Mainframe) Strong understanding of DevOps, CI/CD, and infrastructure as code
Strong strategic thinking with ability to translate business goals into technology outcomes Excellent incident leadership and crisis management skills Proven track record of driving automation and operational transformation Strong stakeholder management and executive communication skills
Expertise in ITIL / Service Management frameworks Strong analytical, problem-solving, and decision-making capabilities Ability to manage high-pressure situations and multiple priorities
DBS is more than a bank - we're shaping the future of finance and communities. With innovation at our core and impact in our DNA, we go beyond banking to build careers, relationships, and a better world. DBS is a leading financial services group in Asia with a presence in 19 markets. Headquartered and listed in Singapore, DBS is in the three key Asian axes of growth: Greater China, Southeast Asia and South Asia. Recognised for its global leadership, DBS has been named “World’s Best Bank” by Global Finance, “World’s Best Bank” by Euromoney and “Global Bank of the Year” by The Banker. The bank is at the forefront of leveraging digital technology to shape the future of banking, having been named “World’s Best Digital Bank” by Euromoney and the world’s “Most Innovative in Digital Banking” by The Banker. In addition, DBS has been accorded the “Safest Bank in Asia” award by Global Finance for 15 consecutive years from 2009 to 2023. DBS provides a full range of services in consumer, SME and corporate banking. As a bank born and bred in Asia, DBS understands the intricacies of doing business in the region’s most dynamic markets and is committed to building lasting relationships with customers. With its extensive network of operations in Asia and emphasis on engaging and empowering its staff, DBS presents exciting career opportunities.